NVIDIA Logo

NVIDIA

Senior Software Engineer, NeMo Core Platform

Posted 2 Days Ago
Be an Early Applicant
In-Office or Remote
5 Locations
184K-357K Annually
Senior level
In-Office or Remote
5 Locations
184K-357K Annually
Senior level
Build and lead development of NVIDIA’s open-source NeMo Core Platform for executing, evaluating, deploying, and optimizing agentic AI systems across local, Docker, Kubernetes, Slurm, on-premises, and air-gapped environments. Design platform APIs, plugins, SDKs, job orchestration, authentication, RBAC, observability, testing, and reliability features. Provide technical leadership through architecture and code reviews, mentoring, and ownership of complex cross-component problems.
The summary above was generated by AI

We are looking for a Senior Software Engineer to help build NeMo Platform, NVIDIA’s product for developing, evaluating, deploying, and operating AI systems at scale. This role is for a senior engineer/architect for our Core team which owns and ships an open source plugin-based AI platform for running and optimizing Agents targeting multiple compute backends (local/docker, Kubernetes, Slurm, etc.).
 

As AI systems become more autonomous and more deeply integrated into real workflows, teams need robust APIs and orchestration systems for running, monitoring, and optimizing Agents at scale. Increasingly, the users of these systems are themselves autonomous or semi-autonomous Agents. The NeMo Platform group is building a sophisticated agent execution framework to enable agents to automatically run hundreds of experiments in parallel to find the most efficient agent architectures for our customers. This is important product engineering research for making agents more sustainable. AI systems are not yet nearly as efficient as they can be, and systems like NeMo Platform will allow large scale AI consumers to automatically tune their agents to use fewer tokens and rely on more efficient models with better throughput.
 

What you'll be doing:

  • Working in a product research environment where we place big bets on where the future is heading, adapting in real time as we build alongside an industry that is constantly evolving with us. This means fast iteration, high ownership, pragmatic decisions, and performance-minded implementation under production constraints

  • Designing an Agentic Execution system that are flexible enough to work in many environments (local, k8s, Slurm, on prem / air-gapped)

  • Provide senior technical leadership through design reviews, code reviews, mentoring, and ownership of ambiguous cross-component problems

  • Building and maintain our Core Platform APIs for running jobs, storing data, entities, secrets, and RBAC and Auth

  • Extending our flexible Plugin Architecture that makes it easy for many teams and external customers to install new capabilities into our system

  • Building in the open in our OSS repo, keeping up the high standards that the open source community demands

  • Shipping code at the speed of light with an unlimited token budget using best in class agentic coding tools

  • Improving reliability, observability, debuggability, and performance across NeMo Platform, SDKs, plugins, jobs, and developer workflows

  • Building strong test coverage across unit, integration, E2E, Docker, and Kubernetes workflows

What we need to see:

  • BS, MS, or equivalent experience in Computer Science, Computer Engineering, or a related technical field

  • 10+ years of professional software engineering experience building production systems

  • Comfort working in a very fast and ambiguous environment

  • Exceptional communication, both verbal and written. This includes the ability to produce and review high quality architectural RFCs, and to discuss them clearly with the right level of technical detail for the right people (Engineer, Product, Marketing, etc.)

  • Strong system design skills, with a pragmatic flexibility and phenomenal instincts to invent robust systems quickly without over-complicating. Strong understanding of reliability, scalability, security, and performance tradeoffs in production infrastructure

  • Experience with distributed systems, cloud-native services, containers, Kubernetes, and job orchestration

  • Excellent Python engineering skills, including API design, typing, testing, debugging, performance analysis, and maintainable software design

  • Experience designing SDKs, libraries, plugins, CLIs, or other developer-facing interfaces

  • Ability to work independently, define technical scope, break down ambiguous problems, and drive work across team boundaries

Ways to stand out from the crowd:

  • Experience building, deploying, and iterating on production agentic AI systems at scale in Kubernetes

  • Experience with sophisticated plugin architectures

  • Strong ability to connect technical evaluation work to business outcomes, product quality, user experience, reliability, or operational efficiency

  • Experience with enterprise AI systems where measurement, regression testing, observability, governance, and continuous improvement are required for production deployment

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 4, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

29 Minutes Ago
Remote or Hybrid
United States
Mid level
Mid level
Artificial Intelligence • Cloud • Payments • Software • Business Intelligence • Generative AI • Automation
Administer hybrid Windows Server infrastructure across colocation, on-premises, Azure, and GCP environments. Manage virtual machines, cloud resources, patching, upgrades, backups, monitoring, alerts, incident response, vulnerability remediation, and root cause analysis. Develop automation using PowerShell, Python, Ansible, and infrastructure-as-code tools for provisioning, remediation, and self-healing workflows. Maintain documentation and runbooks while collaborating with engineering, IT support, security, and operations teams.
Top Skills: AnsibleAppdynamicsArm TemplatesBashDatadogEsxiGitlab Ci/CdGoogle Cloud Platform (Gcp)GrafanaAzureNew RelicOpentelemetryPatchmypcPowershellPythonSccmSignozSolarwindsTerraformVcenterVmware VsphereWindows ServerWsus
An Hour Ago
Remote or Hybrid
Charlotte, NC, USA
45K-85K Annually
Junior
45K-85K Annually
Junior
Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Handles inbound calls and warm leads to assess customers’ insurance needs, recommend appropriate property and casualty products, and convert prospects into policyholders. The role includes paid training and licensing, customer consultation, sales closing, brand representation, and maintaining strong service standards. Representatives work remotely on a fixed weekday and weekend schedule and must maintain a dedicated professional workspace with high-speed wired internet.
Top Skills: High-Speed Wired InternetPc
An Hour Ago
Remote or Hybrid
District of Columbia, USA
161K-247K Annually
Senior level
161K-247K Annually
Senior level
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Leads GM Energy’s customer success and experience strategy across products, channels, partners, and the customer lifecycle. Defines customer journeys, KPIs, scalable engagement and resolution processes, and CX transformation initiatives. Advocates for customer needs across cross-functional teams, oversees launch readiness and platform improvements, manages escalations, develops executive relationships, and builds a high-performing customer success organization.
Top Skills: Ai-Enabled SupportCRMDashboardsOnecrmSelf-Service PlatformsTelephonyVoc/Csat

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account