Fuse Energy Logo

Fuse Energy

AI Inference Engineer

Reposted 10 Days Ago
In-Office or Remote
Hiring Remotely in United States
Mid level
In-Office or Remote
Hiring Remotely in United States
Mid level
Define and build Fuse's inference-serving architecture and stack, including request routing, batching, scheduling, and autoscaling for latency-sensitive workloads. Own model-level optimizations (quantisation, distillation, speculative decoding), translate performance SLAs into capacity plans, and integrate low-level GPU/CUDA performance work with serving infrastructure. Set benchmarks, tooling, and act as technical owner of inference performance and reliability.
The summary above was generated by AI

Fuse Energy is an energy startup on a mission to make energy abundant and affordable, fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system.

We've raised over $200M from top-tier investors including Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, 20VC, Hummingbird and Collaborative Fund, alongside strategic angels including Nico Rosberg and GPs behind Meta, Revolut, Spotify and Uber.

We're building a fully integrated energy company: developing our own solar, batteries and other generation projects, building our own hardware, improving and developing grid infrastructure, trading power in real time, using AI across the business, and installing distributed energy in homes. By selling directly to consumers we cut out the middleman, lower costs and pass the savings on to our customers.

As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch, and we're looking for the founding engineer to own the latter. Reporting directly to the CTO, you'll own the layer above kernels and hardware: how models actually get served, scaled and delivered against committed performance targets. Few companies can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of our offering.

Responsibilities
  • Define Fuse's inference serving strategy and architecture from first principles
  • Design and build the serving stack: request routing, batching, scheduling and autoscaling for high-throughput, latency-sensitive inference workloads
  • Own model-level optimisation strategy for serving, deciding where and how to apply quantisation, distillation, speculative decoding and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers
  • Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)
  • Translate throughput, latency and uptime commitments into concrete technical specifications and serving capacity plans
  • Act as direct technical owner of inference performance and reliability
  • Work closely with the CUDA and GPU engineering teams to integrate custom kernels and hardware performance work cleanly into the serving layer
  • Set the standards, tooling and benchmarks this function will run on as it grows

Requirements
  • 4+ years building or operating large-scale inference serving systems, or equivalent strong project/industry experience
  • Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding)
  • Strong systems thinking, able to reason about the full path from incoming request to served response across a large cluster
  • Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system
  • A track record of making high-stakes architecture calls and owning the outcome
  • Comfort operating without a playbook: this is a founding role shaping a new function, not joining an established one
  • Bonus: Triton or custom ML inference/training frameworks; autoscaling or capacity planning for large-scale inference; multi-tenant serving or SLA-driven infrastructure; background at a hyperscaler, frontier AI lab or large-scale distributed inference system; Kubernetes/Slurm; interest in energy markets, grid systems or sustainability-focused compute

Benefits
  • Competitive salary and eligibility for equity
  • Biannual bonus scheme
  • Fully expensed tech to match your needs
  • Private health insurance
  • Breakfast and dinner allowance for office-based employees

As we hire globally, benefits vary by location.

Similar Jobs

Senior level
Artificial Intelligence • Hardware • Software • Semiconductor
Design, build, and operate CI/CD, Kubernetes-based platforms, deployment automation, and observability for engineering workflows. Improve reliability, performance, and scalability across cloud and on-prem environments, debug cross-boundary failures, perform root-cause analysis, and deliver durable platform software and self-service tooling.
Top Skills: Argo CdArtifact RepositoriesAWSCi/CdContainerized EnvironmentsCustom ResourcesHelmKubernetesKubernetes OperatorsLinuxMtlsPackage RegistriesPythonShellTerraformTls
18 Days Ago
Remote or Hybrid
US
Mid level
Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
Design, implement, and maintain Python frameworks and services that orchestrate distributed engineering workflows across machines and clusters. Build scheduling, execution, resource management, failure recovery, and test infrastructure. Define APIs and abstractions, reason about concurrency and distributed-systems behavior, debug complex multi-system issues, write automated tests and documentation, and partner with platform, CI, release, QA, and product teams to deliver scalable infrastructure.
Top Skills: AsyncioBuild SystemsCi SystemsCluster SchedulersConcurrent.FuturesContainersKubernetesMultiprocessingPytestPythonRelease InfrastructureRemote Execution Systems
18 Days Ago
Remote
US
Mid level
Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
Integrate, validate, and productionize cross-stack inference features across AI frameworks, runtime, compiler, kernels, distributed systems, and hardware. Drive zero-to-one projects, debug system-wide failures, manage accelerated timelines, and improve automation, diagnostics, and repeatable integration practices while collaborating across software and hardware teams.
Top Skills: Ai FrameworksC++Cloud InfrastructureCluster OrchestrationCompilersContainersDistributed SystemsGoHigh-Performance ComputingKernelsLlmsMicroservicesObservabilityPerformance DebuggingProfilingPythonRuntimes

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account