At Netflix, our mission is to entertain the world. Together, we are writing the next episode - pushing the boundaries of storytelling, global fandom and making the unimaginable a reality. We are a dream team obsessed with the uncomfortable excitement of discovering what happens when you merge creativity, intuition and cutting-edge technology. Come be a part of what’s next.
About the RoleAI for Member Systems (AIMS) runs the AI systems behind every recommendation, search result, and personalized experience for 300M+ members. The stack powering it is large and battle tested, built to meet the demands of its time, and remarkably effective at doing so. But AI/ML is moving fast, and the infrastructure that got us here needs to evolve to meet what's next: new model paradigms, tighter cost and efficiency expectations, and the operational maturity that comes with running AI at this scale.
Platform Systems is the engineering foundation of AIMS, owning reliability, scalability, cost efficiency, and developer experience across the org. We are looking for a Staff ML Software Engineer to own the observability, cost, and platform subsystems that support next generation AI workflows and keep the AIMS AI/ML stack trustworthy and ready for what's next, and to contribute to modernizing it. This is a high leverage role that cuts across the org. The work you do here will define how AIMS builds and operates AI/ML systems for the next decade.
ResponsibilitiesDesign, build, and operate subsystems for observability, evaluation, and tooling that Netflix's next generation ML architecture depends on.
Prove the subsystems on AIMS's current operations first, including anomaly detection, root cause analysis, and operational automation.
Design and build observability systems, including observability primitives built for next generation ML systems rather than just classical ML, that give AIMS ML practitioners deep visibility into model behavior, training pipeline health, serving latency, and data quality, making issues detectable and diagnosable before they become incidents.
Identify and drive cost optimization across AIMS training and serving infrastructure, developing frameworks and tooling, increasingly automated, that make compute efficiency a first class concern rather than an afterthought.
Architect reliability improvements across the AIMS AI/ML stack, reducing toil, improving on-call ergonomics, and setting the standard for operational excellence across the org.
Contribute to the target architecture and migration path for the modernized AIMS AI/ML stack, coordinating with the teams driving that effort.
Continuously evaluate emerging infrastructure patterns, model paradigms, and platform capabilities, and translate them into a forward looking roadmap before they become urgent migrations.
Significant experience designing, building, and operating production AI/ML systems at scale, including training pipelines and familiarity with model serving and online inference under high traffic.
Hands-on experience building subsystems that support advanced agentic architectures, such as memory, trace, eval, and replay pipelines, or orchestration and routing layers for complex model systems. This means you've built the orchestration or control logic itself, not just called an API from a script.
Strong software engineering fundamentals with deep Python expertise and working proficiency in at least one JVM language (Scala or Java).
Proven track record of improving AI/ML system reliability, reducing infrastructure costs, and improving operational scalability.
Experience building observability and monitoring systems for AI/ML workloads; you understand what good visibility looks like across training, serving, and data pipelines.
Strong distributed systems background, including batch processing at scale and real time serving infrastructure.
Collaboration with partner teams to drive technical programs across functions, setting direction, managing dependencies, and building consensus without formal authority.
High technical judgment: able to identify common patterns, build reusable frameworks, and make pragmatic calls on what to invest in, what to defer, and what to leave alone.
Comfortable operating without full information; you can scope a problem, define an approach, and adjust course as you learn more.
Familiarity with LLM evaluation, trace, or replay tooling, such as LLM observability platforms or debugging frameworks for complex model systems.
Familiarity with modern AI/ML infrastructure patterns including feature stores, model serving platforms, and experiment frameworks.
Hands on experience migrating production AI/ML systems across technology generations.
Applied experience in personalization domains such as recommendation systems, search, or discovery.
Generally, our compensation structure consists solely of an annual salary; we do not have bonuses. You choose each year how much of your compensation you want in salary versus stock options. To determine your personal top of market compensation, we rely on market indicators and consider your specific job family, background, skills, and experience to determine your compensation in the market range. The range for this role is $600,000.00 - $1,066,000.00. This compensation range will vary based on location.
Netflix provides comprehensive benefits including Health Plans, Mental Health support, a 401(k) Retirement Plan with employer match, Stock Option Program, Disability Programs, Health Savings and Flexible Spending Accounts, Family-forming benefits, and Life and Serious Injury Benefits. We also offer paid leave of absence programs. Full-time hourly employees accrue 35 days annually for paid time off to be used for vacation, holidays, and sick paid time off. Full-time salaried employees are immediately entitled to flexible time off. See more details about our Benefits here.
Netflix is a unique culture and environment. Learn more here.
Inclusion is a Netflix value and we strive to host a meaningful interview experience for all candidates. If you want an accommodation/adjustment for a disability or any other reason during the hiring process, please send a request to your recruiting partner.
We are an equal-opportunity employer and celebrate diversity, recognizing that diversity builds stronger teams. We approach diversity and inclusion seriously and thoughtfully. We do not discriminate on the basis of race, religion, color, ancestry, national origin, caste, sex, sexual orientation, gender, gender identity or expression, age, disability, medical condition, pregnancy, genetic makeup, marital status, or military service.
Job is open for no less than 7 days and will be removed when the position is filled.
Similar Jobs
What you need to know about the Charlotte Tech Scene
Key Facts About Charlotte Tech
- Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
- Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
- Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
- Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
- Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus


