Wizard AI Jobs

Senior Machine Learning Engineer (Inference Platform)

Wizard AI

Senior Machine Learning Engineer (Inference Platform)

Reposted 19 Days Ago

Remote

Hiring Remotely in USA

200K-250K Annually

Senior level

Remote

Hiring Remotely in USA

200K-250K Annually

Senior level

As a Senior MLOps Engineer, you will manage production ML systems, define lifecycle strategies, optimize ML pipelines, and collaborate with cross-functional teams to enhance ML operations.

The summary above was generated by AI

About Wizard AI

At Wizard AI, we’re building the top-performing AI Shopping Agent that delivers the best products from across the web with unmatched accuracy, quality, and trust. Our ML models power the core of our platform, and we’re looking for a Senior Machine Learning Engineer to own how they run in production reliably, efficiently, and at scale.

The Role

As a Senior ML Engineer on our Inference Platform, you’ll own the end-to-end lifecycle of production ML serving systems from model packaging and deployment to monitoring, optimization, and scaling. This is not a traditional MLOps role focused solely on pipelines and tooling. You’ll be responsible for the inference infrastructure powering a live conversational shopping agent, operating multiple specialized serving engines under real-world production load.

You’ll own critical decisions around serving architecture, performance, reliability, and scalability, working closely with ML Engineers, Data teams, Product, and DevOps to ensure models move seamlessly from experimentation into high-performance production systems.

What You'll Do

Own and evolve our multi-engine inference platform, supporting a variety of model types and serving requirements.
Build and improve production ML pipelines — taking models from experimentation to reliable, high-throughput serving.
Define and implement model versioning, rollout, rollback, and lifecycle management strategies that ensure reproducibility and operational reliability.
Define and enforce serving-layer SLAs, including latency, availability, GPU utilization, Time-to-First-Token (TTFT), and Inter-Token Latency (ITL).
Build observability, monitoring, alerting, and operational tooling for production inference systems.
Apply software engineering best practices, including testing, CI/CD integration, and reproducibility across ML workflows.
Optimize inference performance through efficient resource utilization, hardware-aware serving strategies, and cost-conscious infrastructure design.
Ensure ML serving systems are secure, scalable, and operationally resilient.
Partner with ML, Data, Product, and DevOps teams to turn ideas into production systems, driving the technical decisions on serving and scale.

What We're Looking For

Bachelor's or Master's degree in Computer Science, Data Science, Engineering, or a related field, or equivalent practical experience.
5–8+ years of experience in Software Engineering, ML Engineering, Platform Engineering, or Infrastructure Engineering, with direct ownership of production ML serving systems.
Hands-on experience running an LLM serving engine (vLLM, TGI, TensorRT-LLM, or SGLang) in production under real load — not just managed or hosted endpoints.
Strong Python skills and software engineering fundamentals, combined with deep systems and infrastructure knowledge.
Experience with cloud platforms such as AWS, GCP, or Azure, and familiarity with ML lifecycle tooling, experimentation platforms, and model registries.
Strong grasp of inference performance — continuous batching, KV-cache and GPU-memory behavior, quantization, and CPU-versus-GPU bottlenecks — with the instinct to profile before tuning.
Experience serving heterogeneous workloads, including LLMs, embedding models, and extraction models, each with distinct latency, throughput, and scaling requirements.
Demonstrated ability to balance latency, throughput, reliability, and infrastructure cost while operating production-scale ML systems.
Experience in high-growth startup environments and comfort operating in fast-moving, evolving technical landscapes.

What Success Looks LikeReliable, Scalable Inference Systems

Production serving infrastructure operates with clear SLAs, strong observability, and minimal downtime. Latency, availability, throughput, and GPU utilization are actively measured and optimized as platform demands grow.

End-to-End Ownership

You own the complete serving lifecycle — from deployment and release management through monitoring, optimization, and scaling — enabling ML engineers to ship quickly while maintaining reliability and reproducibility.

Technical Leadership and Impact

You shape the future of Wizard's inference platform, driving key architectural decisions that improve performance, reduce infrastructure costs, and support the next generation of AI-powered shopping experiences.

Similar Jobs

PNC Bank

Product Manager

An Hour Ago

Remote or Hybrid

USA

134K-250K Annually

Senior level

134K-250K Annually

Senior level

Machine Learning • Payments • Security • Software • Financial Services

Senior digital product manager for retail lending who defines digital strategy, designs and prioritizes digital experiences, builds business cases, manages development and launches, partners with stakeholders (Technology, Marketing, MIS, LOB), coaches product teams, and supports risk, compliance, and audit needs.

Top Skills: Agile Web DevelopmentBusiness Requirements DocumentationData VisualizationIt ArchitectureJavaScriptWireframing

PNC Bank

Software Engineering Manager

An Hour Ago

Remote or Hybrid

USA

100K-223K Annually

Senior level

100K-223K Annually

Senior level

Machine Learning • Payments • Security • Software • Financial Services

Lead and develop a team of automation engineers, own test automation backlog and frameworks, drive UI/API/performance testing, integrate AI-driven testing practices, collaborate cross-functionally, hire and mentor staff, and contribute hands-on to automation and CI/CD improvements.

Top Skills: Agentic AiApache JmeterAWSAzureBitbucketCi/CdContainerizationCypressGCPGenerative AiGithub ActionsJavaScriptJenkinsMonitoringObservabilityPlaywrightPostmanPrompt EngineeringPythonRest ApisSeleniumTypescript

Federal Reserve Bank of Boston

Security Engineer

2 Hours Ago

Remote

USA

150K-224K Annually

Senior level

150K-224K Annually

Senior level

Fintech • Information Technology • Payments • Sharing Economy • Financial Services • Cryptocurrency

Senior security engineer responsible for developing automation and tooling, building and deploying security solutions, conducting incident response and proactive threat hunting, supporting investigations through data analysis, and advising development teams. Must document solutions, participate in agile teams, mentor juniors, and continuously research evolving security threats.

Top Skills: AWSAws CodedeployAzureBashCircleCICloud IamContainer OrchestrationGithub ActionsGitlab PipelinesGoJavaLog ManagementPythonTravisci

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus