NVIDIA Logo

NVIDIA

Senior Solutions Architect, Generative AI

Posted One Month Ago
Be an Early Applicant
In-Office or Remote
2 Locations
184K-357K Annually
Senior level
In-Office or Remote
2 Locations
184K-357K Annually
Senior level
Architect and optimize large-scale GPU AI infrastructure: profile distributed training/inference, diagnose networking and system bottlenecks, design high-performance clusters, lead POCs, automate benchmarking, and drive technical engagements with customers and internal teams.
The summary above was generated by AI

NVIDIA is looking for an AI Solutions Architect with deep, hands-on experience in large-scale GPU systems. This role involves working with some of the world’s leading consumer internet companies and frontier labs building foundation models. Primary responsibilities include accelerating customer workloads, designing high-performance AI infrastructure, and leading technical engagements around NVIDIA technologies. We work with the world’s most successful technology companies, uniquely positioning you to observe and influence emerging infrastructure trends using the latest advancements. Join us in this exciting endeavor!

What You’ll Be Doing:

  • Collaborating closely with customers to maximize GPU utilization and end-to-end workload throughput while improving infrastructure reliability and reducing infrastructure costs.

  • Designing and optimizing large-scale AI clusters across GPU compute, high-performance networking, storage, workload scheduling, orchestration, and observability.

  • Profiling distributed training and inference workloads to identify bottlenecks across GPUs, CPUs, memory, network fabrics, storage systems, and software stack.

  • Diagnosing complex infrastructure and distributed systems issues spanning InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch.

  • Leading proof-of-concepts and performance studies for large-scale AI infrastructure, developing benchmarking tools, automation, runbooks, and technical collateral as needed.

  • Partnering with NVIDIA’s engineering, product, and sales teams to secure design wins and drive innovative solutions based on customer requirements and field feedback.

What We Need To See:

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or another Engineering field, or equivalent experience.

  • 6+ years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a related technical role.

  • Deep understanding of Linux systems, distributed computing, GPU architectures, and the hardware and software components of large-scale AI clusters.

  • Hands-on experience designing, deploying, operating, or troubleshooting high-performance GPU networks in on-premises or cloud environments using technologies such as InfiniBand, RoCE, or GPUDirect RDMA.

  • Experience debugging NCCL communication and distributed collective performance, including topology, transport, congestion, routing, and host-level configuration issues.

  • Experience profiling AI workloads and identifying performance bottlenecks across compute, networking, storage, and orchestration layers.

  • Experience with cluster schedulers and orchestration platforms such as Kubernetes and Slurm, along with containers and production monitoring systems.

  • Proficiency with Python, shell scripting, or similar languages for infrastructure automation, benchmarking, and systems troubleshooting.

Ways To Stand Out From The Crowd:

  • Experience architecting and operating large-scale production GPU clusters for distributed training or inference.

  • Deep expertise with NVIDIA infrastructure technologies such as DGX/HGX systems, NVLink, NVSwitch, NCCL, InfiniBand, and Spectrum-X.

  • Hands-on experience using tools and telemetry such as NCCL tests, DCGM, Nsight Systems, fabric counters, and host- or switch-level diagnostics to isolate performance and reliability issues.

  • Understanding of network topology, congestion control, collective communication patterns, and their impact on distributed AI workload performance.

  • Experience optimizing storage and data pipelines to sustain high-throughput training and inference workloads.

We make extensive use of conferencing tools, but occasional travel (20%) is required for local on-site visits to customers and conferences. We are open to remote work. We look forward to having you join our team!

With competitive salaries and a generous benefits package, NVIDIA is recognized as one of the technology world’s most sought-after employers. This role offers a chance to make a broad impact at NVIDIA by advancing innovation with our consumer internet & frontier labs partners. Are you inventive, diligent, committed, and driven? Do you enjoy tackling challenges? If so, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 3, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

14 Days Ago
In-Office or Remote
4 Locations
184K-288K Annually
Senior level
184K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Partners with customers and internal teams to design, deploy, and optimize generative AI and LLM inference solutions. Develops proofs of concept, supports NVIDIA technology adoption, analyzes AI workload performance and power efficiency on Kubernetes, and advises developers, researchers, and data scientists. The role requires expertise in deep learning frameworks, GPU orchestration, MLOps, containerization, observability, and LLM inference, with occasional travel to conferences and customers.
Top Skills: CC++ContainerizationKubernetesMlopsMonitoringMulti-Instance Gpu (Mig)Nvidia DynamoNvidia GpusNvidia NimObservabilityPythonPyTorchTensorFlowTensorrtTensorrt-Llm
20 Days Ago
Remote
US
180K-225K Annually
Senior level
180K-225K Annually
Senior level
Security • Cybersecurity
Lead design, prototyping, and production deployment of Generative AI/LLM solutions and the COE reference architecture (RAG, agents, model routing, evaluation, guardrails). Evaluate vendors, integrate GenAI into enterprise systems, mentor technical leads, and drive reusable patterns and responsible AI practices across hybrid cloud environments.
Top Skills: Agent FrameworksAgent OrchestrationAWSAws BedrockAzure (Commercial And Gov)Azure Ai FoundryAzure Openai ServiceEvaluation HarnessesGCPLangchainLlamaindexModel RoutingPyTorchRagSagemakerSemantic KernelVector Databases
28 Days Ago
In-Office or Remote
4 Locations
184K-288K Annually
Senior level
184K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
As a Senior Solutions Architect, you will assist customers in building solutions with NVIDIA's AI technology, focusing on Generative AI and Large Language Models while collaborating across teams for performance analysis.
Top Skills: AIDeep LearningDockerDynamoGpuKubernetesLlmNvidia NimPythonPyTorchTensorFlowTensorrt

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account