Fluidstack Logo

Fluidstack

Infrastructure Engineer (Compute)

Reposted 15 Days Ago
Remote
29 Locations
Senior level
Remote
29 Locations
Senior level
Design, deploy, and manage compute infrastructure for GPU clusters, ensuring performance and reliability while automating server lifecycle tasks.
The summary above was generated by AI
About FluidStack

Fluidstack is the AI Cloud Platform. We build GPU supercomputers for top AI labs, governments, and enterprises. Our customers include Mistral, Poolside, Black Forest Labs, Meta, and more.

Our team is small, highly motivated, and focused on providing a world class supercomputing experience. We put our customers first in everything we do, working hard to not just win the sale, but to win repeated business and customer referrals.

We hold ourselves and each other to high standards. We expect you to care deeply about the work you do, the products you build, and the experience our customers have in every interaction with us.

You must work hard, take ownership from inception to delivery, and approach every problem with an open mind and a positive attitude. We value effectiveness, competence, and a growth mindset.

About the Role

We are looking for an Infrastructure Engineer (Compute) to design, deploy, and manage the compute infrastructure powering Fluidstack's GPU clusters. You will be responsible for ensuring the performance, scalability, and reliability of our compute resources, working closely with hardware and software teams to support our AI workloads.

Focus
  • Design and implement GPU/ASIC infrastructure at the server, rack, and system level.

  • Troubleshoot complex GPU and compute system related failures.

  • Develop and maintain hardware/firmware management services.

  • Automate all aspects of the server lifecycle.

  • Own end-to-end compute lifecycle, including partnering with vendors on RMAs.

  • Serve as the main point of contact for hardware escalation and troubleshooting.

  • Monitor system performance, identifying and resolving bottlenecks.

  • Automate deployment and management tasks to improve efficiency.

  • Collaborate with storage and network teams to ensure cohesive infrastructure operations.

About You
  • 5+ years of experience in compute infrastructure engineering.

  • Strong knowledge of Linux systems administration and performance tuning.

  • Experience with bare metal provisioning tools (MaaS, Metal3, Tinkerbell, or other).

  • Familiarity with GPU hardware and workload optimization, especially kernel and driver level requirements.

  • Proficiency in automation tools (e.g., Ansible, Terraform).

  • Experience operating Kubernetes and SLURM clusters.

Benefits
  • Competitive total compensation package (salary + equity).

  • Retirement or pension plan, in line with local norms.

  • Health, dental, and vision insurance.

  • Generous PTO policy, in line with local norms.

  • Fluidstack is remote first, but has offices in key hubs. For all other locations, we provide access to WeWork.

Top Skills

Ansible
Kubernetes
Linux
Maas
Metal3
Slurm
Terraform
Tinkerbell

Similar Jobs

17 Hours Ago
Remote or Hybrid
8 Locations
Senior level
Senior level
Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
As a Senior Incident Response Analyst, you'll investigate security incidents, enhance security measures, coach analysts, and collaborate with teams to strengthen incident response capabilities.
Top Skills: Cloud Computing ServicesDatabaseDlpEmail SecurityEndpoint SecurityFirewallsIamIds/IpsMitre Att&Ck FrameworkProxiesScriptingSIEMSoarWafWeb Content Filtering
17 Hours Ago
Easy Apply
Remote
28 Locations
Easy Apply
Mid level
Mid level
Artificial Intelligence • Machine Learning • Natural Language Processing • Conversational AI
As a DevOps Engineer, you will manage, troubleshoot, and enhance production environments, ensure CI/CD processes, and implement automated solutions.
Top Skills: Amazon DynamodbAnsibleArgocdAWSAws CloudformationAzure Cosmos DbChefContainerdCouchbaseCrossplaneDatadogDockerElk StackGitlab CiGoogle Cloud PlatformGrafanaHelmJenkinsKubernetesLokiAzureMongoDBNomadOctopus DeployOpenshiftOracle Cloud InfrastructurePrometheusPulumiPuppetSaltstackTerraformVictoriametricsZabbix
Yesterday
Remote or Hybrid
10 Locations
Senior level
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
The Data Quality Lead will develop data quality practices, modernize data services, oversee cross-functional teams, and establish analytics capabilities in a digital environment.
Top Skills: AIJavaMachine LearningPythonRSASScalaSnowflakeSparkSQLTableau

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account