NVIDIA Logo

NVIDIA

Senior Datacenter System Software Architect - DGX Cloud

Posted 18 Days Ago
Be an Early Applicant
In-Office or Remote
2 Locations
184K-357K
Senior level
In-Office or Remote
2 Locations
184K-357K
Senior level
Lead the architecture, design, and implementation of DGX cloud clusters. Oversee technical activities, infrastructure workflows, and ensure software integration for AI applications.
The summary above was generated by AI

NVIDIA is hiring engineers to scale up its AI Infrastructure. We expect you to have a strong programming background, a deep understanding of distributed systems, familiarity with software testing and deployment, and excellent communication and planning abilities. We also welcome out-of-the-box thinkers who can provide new ideas with strong at execution bias. Expect to be constantly challenged, improving, and evolving for the better. You and other engineers in this team will help advance NVIDIA's capacity to build and deploy leading infrastructure solutions for a broad range of AI-based applications that affect core data science. What are you waiting for if you're creative, passionate about what you do, and love having fun apply today!

We’re looking for a highly motivated, creative engineer with strong experience in system software to join the DGX Cloud Software Team. You will lead the architecture, design and implementation of our next generation DGX cloud clusters using latest technologies. On this team, you will do full stack deployment including hardware architecture, workload orchestration and application performance tuning. Are you ready to change the next generation of computing? Join us at the forefront of technological advancement.

What you’ll be doing:

  • Lead technical activities for data centers with focus on hybrid deployments between cloud and on-prem

  • Providing expertise in infrastructure workflows, including hardware, software release, workload orchestration and application tuning

  • Provide fast and creative solutions for complex problems and write effective, clear and reliable architecture specification

  • Translate requirements to vision, architecture and roadmap

  • Work with engineering teams across NVIDIA to ensure your software integrates seamlessly from the hardware all the way up to the AI training applications.

What we need to see:

  • Masters or PhD in Computer Science, Computer Engineering, Physics or equivalent experience.

  • 9+ years of experience in this field.

  • Data Sciences, Deep Learning, or Machine Learning coursework

  • Ability to seamlessly shift between Linux system environments to Python programming

  • Programming skills in 1 or more high-level languages (C, C++, Go, Rust, etc)

  • System-level experience with both hardware and software

  • Motivated self-starter with an equal balance of strong problem-solving skills and customer-facing communication skills

  • Strong design, coding, analytical, debugging and problem-solving skills

  • Passion for continuous learning and knowledge transfer. Ability to work concurrently with multiple groups locally and abroad in the organization

Ways to stand out from the crowd:

  • Experience with GPU deep learning and data sciences. Experience using TensorFlow, PyTorch or other DL framework. Experience working with Docker containers, Slurm, Terraform and Kubernetes

  • CUDA programming and NCCL experience. HPC programming experience including MPI, OpenACC, or other parallel programming tools. Hands-on experience with DGX Cloud, NVIDIA AI Enterprise AI Software, Base Command Manager, NEMO and NVIDIA Inference Microservices.

  • Interest in crafting, analyzing and fixing large-scale distributed systems.

  • Systematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 2, 2025.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Top Skills

C
C++
Cuda
Docker
Go
Kubernetes
Mpi
Nccl
Openacc
Python
PyTorch
Rust
Slurm
TensorFlow
Terraform

Similar Jobs

52 Minutes Ago
Remote or Hybrid
Kirkland, WA, USA
141K-239K Annually
Senior level
141K-239K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
The Senior Software Engineer will build scalable code, collaborate with product owners, implement software features, and mentor teammates. Experience with AI integration and testing techniques is required.
Top Skills: C#EclipseGitJavaJavaScriptJenkinsMaven
52 Minutes Ago
Remote or Hybrid
Waltham, MA, USA
131K-230K Annually
Expert/Leader
131K-230K Annually
Expert/Leader
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Manage long-term customer success for a portfolio of CPQ clients. Guide onboarding, adoption, and integration strategies while ensuring measurable ROI and driving AI innovation within customer workflows.
Top Skills: AIAPIsAutomationCpqCRMEcommerceIntegrationsSaaS
52 Minutes Ago
Remote or Hybrid
New York, NY, USA
218K-381K Annually
Expert/Leader
218K-381K Annually
Expert/Leader
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Lead the enterprise OKR program by driving strategic alignment and managing a high-performing team to implement measurement systems across ServiceNow.
Top Skills: Ai-Powered ToolsOkr FrameworkPerformance Management Systems

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account