NextSilicon Logo

NextSilicon

HPC Systems Administrator

Reposted 28 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
Senior level
Remote
Hiring Remotely in United States
Senior level
Administer, provision, and maintain HPC/AI compute, storage, networking, and software stacks. Develop automation for provisioning, configuration management, and monitoring. Install, configure, and optimize job schedulers (e.g., Slurm), deploy MPI and containerized HPC applications, perform benchmarking, tuning, capacity planning, security patching, troubleshooting, and vendor coordination while supporting researchers and documenting procedures.
The summary above was generated by AI
Description

NextSilicon is revolutionizing high-performance computing. Our innovative coprocessor technology dramatically accelerates supercomputers, propelling them into a new era. Our software-defined hardware architecture empowers HPC/AI to deliver groundbreaking discoveries across all areas of advanced research. We're seeking a dynamic and results-oriented HPC/AI Systems Administrator to join our team.

At NextSilicon, everything we do is guided by three core values:

  • Professionalism: We strive for exceptional results through professionalism and unwavering dedication to quality and performance. 
  • Unity: Collaboration is key to success. That's why we foster a work environment where every employee can feel valued and heard. 
  • Impact: We're passionate about developing technologies that make a meaningful impact on industries, communities, and individuals worldwide.

Join our Field Deployment & Systems team as an HPC/AI Systems Administrator.

As an HPC/AI Systems Administrator at NextSilicon, you will be central to sustaining the successful operation of HPC/AI systems. You will stand-up and maintain HPC/AI hardware and software resources. You will tune and configure systems for high-quality benchmarking efforts. You will ensure that the health and accessibility of the HPC/AI systems is top-notch via cluster management tools and capacity planning efforts.

This is a highly technical, execution-focused individual contributor role with no people management or leadership responsibilities at this time.

Location: Hybrid in either our Austin, TX or Minneapolis, MN offices preferred but Remote considered for exceptional candidates.

Requirements
  • Bachelor’s degree in engineering, mathematics, computer science, related field, or equivalent experience. Advanced degree is a plus.
  • 5-10+ years of experience with HPC/AI system administration.
  • Deep understanding of HPC & AI technologies and software ecosystems
  • Hands-on experience configuring, maintaining, and troubleshooting Slurm in large-scale HPC/AI environments.
  • Experience in a fast-paced, entrepreneurial environment is a plus
  • Ability to travel within the USA approx. 4 times per year
  • US citizenship with eligibility to visit US government research facilities
Responsibilities
  • Administer, install, monitor, and maintain HPC/AI systems, including compute nodes, storage, networking, and software stacks.
  • Develop and maintain automation tools for system provisioning, configuration management, and monitoring. 
  • Install, configure, and optimize job scheduling and resource management tools (e.g., Slurm). 
  • Assist in system security, patch management, and troubleshooting operational issues. 
  • Contribute to performance benchmarking, system tuning, and capacity planning. 
  • Deploy and maintain commonly used HPC/AI applications, software stacks, and technologies (e.g., MPI, containers, spack, modules)
  • Document system administration procedures and contribute to knowledge-sharing initiatives.
  • Support researchers by providing technical expertise and resolving escalated support tickets.
  • Participate in vendor coordination, system procurement, and hardware/software lifecycle management.

Similar Jobs

One Month Ago
Remote
United States
98K-132K Annually
Senior level
98K-132K Annually
Senior level
Aerospace • Information Technology • Professional Services • Security • Software
Lead and sustain large-scale HPC systems supporting NOAA/NWS forecasting. Manage scheduler/software stack (PBS Pro/Slurm), optimize performance and throughput, troubleshoot multi-node Linux clusters, develop scripts, and support 24/7 operations with occasional travel and on-call duties.
Top Skills: BashHigh Performance Computing (Hpc)MpiPbs ProPerlPythonRed Hat Enterprise Linux (Rhel)Rocky LinuxSlurmSuse Linux Enterprise Server (Sles)
9 Minutes Ago
In-Office or Remote
138K-172K Annually
Senior level
138K-172K Annually
Senior level
Software
Manage customer renewals, retention, adoption, and expansion for existing SaaS accounts. Build strategic relationships with decision-makers, negotiate contracts, mitigate churn, identify upsell opportunities, and collaborate with account, product, sales, support, legal, and operations teams. Maintain accurate Salesforce records, analyze customer and pipeline data, conduct business reviews, support onboarding, and communicate business value in a quota-carrying environment.
Top Skills: Application Performance MonitoringDevOpsGoogle SuiteExcelMicrosoft PowerpointObservabilitySalesforceSalesloft
11 Minutes Ago
Remote or Hybrid
70K-95K Annually
Entry level
70K-95K Annually
Entry level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Conduct OSINT investigations into Latin American eCrime actors, tools, techniques, and online communities using Spanish and Portuguese language skills. Collect actionable intelligence from social media, forums, hidden services, and dark web environments; identify emerging cyber threats; collaborate on analytical reports; and apply secure operational tradecraft. The role requires strong attention to detail, independent execution, teamwork, adaptability, and the ability to use AI technologies to improve investigative workflows.
Top Skills: Ai TechnologiesDark WebDeep WebHidden ServicesOnline ForumsOsintSocial MediaVirtual Humint

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account