NextSilicon Logo

NextSilicon

HPC Systems Administrator

Reposted 28 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
Senior level
Remote
Hiring Remotely in United States
Senior level
Administer, provision, and maintain HPC/AI compute, storage, networking, and software stacks. Develop automation for provisioning, configuration management, and monitoring. Install, configure, and optimize job schedulers (e.g., Slurm), deploy MPI and containerized HPC applications, perform benchmarking, tuning, capacity planning, security patching, troubleshooting, and vendor coordination while supporting researchers and documenting procedures.
The summary above was generated by AI
Description

NextSilicon is revolutionizing high-performance computing. Our innovative coprocessor technology dramatically accelerates supercomputers, propelling them into a new era. Our software-defined hardware architecture empowers HPC/AI to deliver groundbreaking discoveries across all areas of advanced research. We're seeking a dynamic and results-oriented HPC/AI Systems Administrator to join our team.

At NextSilicon, everything we do is guided by three core values:

  • Professionalism: We strive for exceptional results through professionalism and unwavering dedication to quality and performance. 
  • Unity: Collaboration is key to success. That's why we foster a work environment where every employee can feel valued and heard. 
  • Impact: We're passionate about developing technologies that make a meaningful impact on industries, communities, and individuals worldwide.

Join our Field Deployment & Systems team as an HPC/AI Systems Administrator.

As an HPC/AI Systems Administrator at NextSilicon, you will be central to sustaining the successful operation of HPC/AI systems. You will stand-up and maintain HPC/AI hardware and software resources. You will tune and configure systems for high-quality benchmarking efforts. You will ensure that the health and accessibility of the HPC/AI systems is top-notch via cluster management tools and capacity planning efforts.

This is a highly technical, execution-focused individual contributor role with no people management or leadership responsibilities at this time.

Location: Hybrid in either our Austin, TX or Minneapolis, MN offices preferred but Remote considered for exceptional candidates.

Requirements
  • Bachelor’s degree in engineering, mathematics, computer science, related field, or equivalent experience. Advanced degree is a plus.
  • 5-10+ years of experience with HPC/AI system administration.
  • Deep understanding of HPC & AI technologies and software ecosystems
  • Hands-on experience configuring, maintaining, and troubleshooting Slurm in large-scale HPC/AI environments.
  • Experience in a fast-paced, entrepreneurial environment is a plus
  • Ability to travel within the USA approx. 4 times per year
  • US citizenship with eligibility to visit US government research facilities
Responsibilities
  • Administer, install, monitor, and maintain HPC/AI systems, including compute nodes, storage, networking, and software stacks.
  • Develop and maintain automation tools for system provisioning, configuration management, and monitoring. 
  • Install, configure, and optimize job scheduling and resource management tools (e.g., Slurm). 
  • Assist in system security, patch management, and troubleshooting operational issues. 
  • Contribute to performance benchmarking, system tuning, and capacity planning. 
  • Deploy and maintain commonly used HPC/AI applications, software stacks, and technologies (e.g., MPI, containers, spack, modules)
  • Document system administration procedures and contribute to knowledge-sharing initiatives.
  • Support researchers by providing technical expertise and resolving escalated support tickets.
  • Participate in vendor coordination, system procurement, and hardware/software lifecycle management.

Similar Jobs

One Month Ago
Remote
United States
98K-132K Annually
Senior level
98K-132K Annually
Senior level
Aerospace • Information Technology • Professional Services • Security • Software
Lead and sustain large-scale HPC systems supporting NOAA/NWS forecasting. Manage scheduler/software stack (PBS Pro/Slurm), optimize performance and throughput, troubleshoot multi-node Linux clusters, develop scripts, and support 24/7 operations with occasional travel and on-call duties.
Top Skills: BashHigh Performance Computing (Hpc)MpiPbs ProPerlPythonRed Hat Enterprise Linux (Rhel)Rocky LinuxSlurmSuse Linux Enterprise Server (Sles)
10 Minutes Ago
Remote
United States
145K-145K Annually
Senior level
145K-145K Annually
Senior level
Security • Cybersecurity
Design, deploy, and maintain Google Cloud infrastructure supporting internal IT. Build repeatable Terraform-based deployments, manage GCP architecture and networking, implement security and compliance controls, monitor performance and costs, troubleshoot cloud environments, and maintain documentation. Collaborate with IT, security, and engineering teams while supporting Azure, AWS, on-premises connectivity, Kubernetes workloads, and infrastructure integration during acquisitions.
Top Skills: AWSBashClaude CodeCloud InterconnectCloud MonitoringCloud StorageCompute EngineDatadogDnsEntra IdFirewallsGoogle Cloud Platform (Gcp)Google Kubernetes Engine (Gke)IamInfrastructure As Code (Iac)KubernetesLoad BalancingAzureNistPythonRoutingSecurity Command CenterSoc 2SwitchingTcp/IpTerraformVlansVpcVpc Service ControlsVpn
19 Minutes Ago
Remote or Hybrid
114K-165K Annually
Senior level
114K-165K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Drive new business revenue through SaaS sales by planning accounts and territories, researching prospects, and executing field-based sales activities. Build C-suite relationships, coordinate account strategies across cross-functional teams, advise customers on IT roadmaps, and engage specialists to support deals. The role requires selling cybersecurity solutions, developing new business, negotiating agreements, achieving sales targets, and serving customers in the energy and transportation industries.
Top Skills: AICybersecurityIdentity SecuritySaaSServicenowVeza

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account