BJAK

Principal Machine Learning Engineer

Posted 2 Days Ago

Remote

Hiring Remotely in United States

Expert/Leader

Remote

Hiring Remotely in United States

Expert/Leader

The Principal Machine Learning Engineer will build and maintain end-to-end ML systems, fine-tune models, and manage production deployments, ensuring robustness and efficiency.

The summary above was generated by AI

About the Role

A1 is building a proactive AI system that understands context across conversations, plans actions, and carries work forward over time.

You will be responsible for turning research direction into working, production-grade ML systems. This role owns the execution layer of A1’s intelligence – training pipelines, inference systems, evaluation tooling, and deployment.

Focus

Build and own end-to-end ML pipelines spanning data, training, evaluation, inference, and deployment.
Fine-tune and adapt models using state-of-the-art methods such as LoRA, QLoRA, SFT, DPO, and distillation.
Architect and operate scalable inference systems, balancing latency, cost, and reliability.
Design and maintain data systems for high-quality synthetic and real-world training data.
Implement evaluation pipelines covering performance, robustness, safety, and bias, in partnership with research leadership.
Own production deployment, including GPU optimization, memory efficiency, latency reduction, and scaling policies.
Collaborate closely with application engineering to integrate ML systems cleanly into backend, mobile, and desktop products.
Make pragmatic trade-offs and ship improvements quickly, learning from real usage.
Work under real production constraints: latency, cost, reliability, and safety

Requirements

Strong background in deep learning and transformer-based architectures.
Hands-on experience training, fine-tuning, or deploying large-scale ML models in production.
Proficiency with at least one modern ML framework (e.g. PyTorch, JAX), and ability to learn others quickly.
Experience with distributed training and inference frameworks (e.g. DeepSpeed, FSDP, Megatron, ZeRO, Ray).
Strong software engineering fundamentals – you write robust, maintainable, production-grade systems.
Experience with GPU optimization, including memory efficiency, quantization, and mixed precision.
Comfort owning ambiguous, zero-to-one ML systems end-to-end.
A bias toward shipping, learning fast, and improving systems through iteration.

Ideal Experience

Experience with LLM inference frameworks such as vLLM, TensorRT-LLM, or FasterTransformer.
Contributions to open-source ML or systems libraries.
Background in scientific computing, compilers, or GPU kernels.
Experience with RLHF pipelines (PPO, DPO, ORPO).
Experience training or deploying multimodal or diffusion models.
Experience with large-scale data processing (Apache Arrow, Spark, Ray).

How We Work

Our organization is very flat and our team is small, highly motivated, and focused on engineering and product excellence. All members are expected to be hands-on and to contribute directly to the company’s mission.

Interview process

If there appears to be a fit, we'll reach to schedule 3, but no more than 4 interviews.

Applications are evaluated by our technical team members. Interviews will be conducted via virtual meetings and/or onsite.

We value transparency and efficiency, so expect a prompt decision. If you've demonstrated the exceptional skills and mindset we're looking for, we'll extend an offer to join us. This isn't just a job offer; it's an invitation to be part of a team that's bringing AI to have practical benefits to billions globally.

Top Skills

Apache Arrow

Deepspeed

Fsdp

Jax

Megatron

PyTorch

Ray

Spark

Zero

Similar Jobs

Atlassian

Principal Machine Learning System Engineer

15 Hours Ago

In-Office or Remote

San Francisco, CA, USA

194K-303K Annually

Senior level

194K-303K Annually

Senior level

Cloud • Information Technology • Productivity • Security • Software • App development • Automation

Lead the construction and maintenance of core ML infrastructure, enabling teams to develop, deploy, and operate ML models effectively. Collaborate with cross-functional teams and drive technical projects from design to launch.

Top Skills: Amazon Web ServicesSparkJavaKotlinPython

Atlassian

Machine Learning Engineer

15 Hours Ago

In-Office or Remote

San Francisco, CA, USA

194K-303K Annually

Senior level

194K-303K Annually

Senior level

Cloud • Information Technology • Productivity • Security • Software • App development • Automation

As a Principal Machine Learning Engineer, you will develop and implement advanced ML algorithms, collaborate across teams, and mentor junior engineers, driving AI integration in Atlassian products.

Top Skills: AWSDatabricksJavaPythonSparkSQL

ServiceNow

Principal Software Engineer

2 Days Ago

Remote or Hybrid

Santa Clara, CA, USA

218K-381K Annually

Expert/Leader

218K-381K Annually

Expert/Leader

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation

Design and develop AI conversational agents and intelligent search systems. Collaborate on scalable AI solutions, mentor team members, and enhance existing products.

Top Skills: AjaxCSSHTMLJavaJavaScriptJSONJunitRestSeleniumTestngXML

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus