BJAK Logo

BJAK

Machine Learning Platform Engineer

Posted One Month Ago
Remote
Hiring Remotely in United States
Mid level
Remote
Hiring Remotely in United States
Mid level
Design, build, and operate ML infrastructure for training, evaluation, deployment, and inference. Improve reliability, scalability, latency, throughput, and cost. Create pipelines, observability, benchmarking, and tooling to enable fast experimentation and productionization of models while diagnosing regressions and bottlenecks.
The summary above was generated by AI

About ActAI

There are over 5 billion users using basic applications today such email, notes, tasks, calendar and they're not AI-native. Our mission is to build proactive applications for anyone in the world, who are not used to complex prompting. We aim to bring intelligence to conversations, errands, organising and workflows, with minimal to no prompting.

Our product focuses on achieving high reliability for long-running workflows, persistent context, and real-world task completion. We believe products will greatly reduce hallucinations.

Our objective is to organise anyone's life, allowing us all to spend time on valuable and meaningful things.

About the Role

As an ML Platform Engineer, you will build the infrastructure and systems that power ActAI's AI capabilities.

You will design and operate the systems behind the AI stack, from model training and evaluation to deployment, inference, observability, and continuous improvement.

You will work closely with AI engineers, researchers, and product engineers to turn models into reliable, scalable, and cost-efficient production systems. You will build the platforms, tooling, and infrastructure that enable the team to experiment quickly and bring AI capabilities to production with confidence.

Focus

  • Build and operate the ML infrastructure and platforms powering A1’s AI products

  • Design systems for model training, evaluation, deployment, inference, and experimentation

  • Build and optimise model serving and inference infrastructure for high-throughput and low-latency workloads

  • Improve reliability, scalability, latency, and cost efficiency of AI systems

  • Develop reliable pipelines for data preparation, training, evaluation, model release, and continuous improvement

  • Build platforms and tooling that enable AI engineers and researchers to experiment, evaluate, and ship models faster

  • Develop evaluation and benchmarking infrastructure to measure model quality, performance, and regressions

  • Build production observability, monitoring, tracing, and alerting for AI/ML workloads

  • Improve AI systems across reliability, scalability, latency, throughput, and cost

  • Identify bottlenecks across the ML stack and continuously improve system performance

  • Work closely with AI engineers, researchers, and product teams to turn evolving model requirements into production-ready infrastructure

Tech Stack

  • Python

  • PyTorch / JAX

  • LLM and ML serving infrastructure such as vLLM, SGLang, or TensorRT-LLM

  • Cloud infrastructure

  • Distributed systems

  • ML/data pipelines and workflow orchestration

  • GPU infrastructure and performance tooling

  • Vector databases and retrieval infrastructure

Ideal Experience

  • Strong software engineering fundamentals and experience building production systems

  • Experience building ML infrastructure, platforms, or production machine learning systems

  • Experience with model deployment, inference, evaluation, or data pipelines

  • Strong understanding of distributed systems and system reliability

  • Ability to write clean, maintainable, production-quality code

  • Comfortable working in ambiguous, fast-moving environments

  • Bias toward ownership, experimentation, and continuous improvement

Outcomes

  • AI infrastructure reliably supports production workloads at scale

  • Models can be trained, evaluated, deployed, and improved efficiently

  • Inference systems deliver strong latency, throughput, reliability, and cost efficiency

  • ML pipelines are reproducible, observable, maintainable, and robust

  • Model and infrastructure regressions are detected quickly and diagnosed efficiently

  • Common ML infrastructure capabilities become reusable platform primitives rather than being rebuilt for every AI product

  • The AI stack can evolve rapidly as new models, architectures, and inference techniques emerge

Similar Jobs

4 Days Ago
Remote or Hybrid
219K-335K Annually
Senior level
219K-335K Annually
Senior level
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Lead development of scalable, reliable continuous integration infrastructure supporting autonomous-vehicle development, machine learning training, simulation, and remote builds. Specialize in Remote Build Execution and a FUSE-based file system, enabling source-code editing and developer workflows. Design and implement productivity improvements, evaluate technologies, influence technical roadmaps, establish engineering best practices, manage technical debt, and mentor engineers while balancing business and customer priorities.
Top Skills: DockerFuseGoGoogle Cloud Platform (Gcp)KubernetesNetworkingPythonRemote Build Execution (Rbe)SshUnix/Linux
11 Days Ago
Remote
United States
145K-250K Annually
Senior level
145K-250K Annually
Senior level
Cloud • Software
Build and operate a secure Kubernetes-based AI/ML platform for government test and evaluation teams. Responsibilities include GPU scheduling, workload orchestration, infrastructure-as-code, GitOps, observability, capacity planning, upgrades, security hardening, compliance support, and operational documentation. The role involves restricted and disconnected environments, automated security tooling, collaboration with government and engineering stakeholders, technical leadership, and mentoring.
Top Skills: Argo CdArtifact SigningAWSAzureCi/CdContainer ScanningDastGitopsGoGCPGpu SchedulingGrafanaHelmKserveKubernetesKubernetes OperatorsLlm-DMlopsNebariNist 800-171Nist 800-53OpentelemetryOpentofuPolicy EnforcementPrometheusPythonRisk Management FrameworkSastTerraformVllm
14 Days Ago
Remote
United States
189K-265K Annually
Expert/Leader
189K-265K Annually
Expert/Leader
Automotive • Insurance • Machine Learning • Mobile • Software
Lead the architecture and development of Root’s machine learning platform for insurance pricing. Build infrastructure for feature pipelines and stores, reproducible training and orchestration, model registries, serving, validation, observability, and research-to-production workflows. Establish versioned contracts and technical standards, automate workflows using LLMs, write and review critical code, influence cross-team strategy, and mentor senior engineers. Collaborate closely with data scientists and researchers to create scalable, reliable, and reproducible ML capabilities.
Top Skills: Distributed SystemsFeature StoresLarge Language Models (Llms)Machine LearningModel ServingPython

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account