Mirantis Logo

Mirantis

Principal HPC Network Engineer (remote in the US)

Posted 12 Days Ago
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Design, deploy, manage, and troubleshoot high-performance InfiniBand and Ethernet networks for HPC. Perform performance tuning, capacity planning, and monitoring. Implement Fortinet security, resolve routing/switching/latency issues, collaborate with compute and storage teams, document architectures, and participate in on-call escalation and upgrades.
The summary above was generated by AI
Company Description

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.  https://www.mirantis.com/

Job Description

Role Overview:
We are seeking a highly skilled Senior HPC Networking Engineer to design, deploy, manage, and troubleshoot high-performance networking environments. The ideal candidate will have deep expertise in InfiniBand technologies, strong general networking knowledge, and hands-on experience with Fortinet solutions. You will play a critical role in ensuring the performance, reliability, and scalability of HPC infrastructure.

Key Responsibilities:

  • Design, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics.

  • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.

  • Manage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations.

  • Perform performance tuning, monitoring, and capacity planning for HPC networking systems.

  • Implement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).

  • Diagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments.

  • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.

  • Develop and maintain documentation for network architecture, configurations, and operational procedures.

  • Participate in on-call rotations and provide escalation support for critical incidents.

  • Lead or contribute to network upgrades, migrations, and new deployments.

Qualifications

Required:

  • 5+ years of experience in network engineering, with a focus on HPC or data center environments.

  • Strong hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA).

  • Solid understanding of networking fundamentals: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.

  • Proven experience deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).

  • Experience with network performance analysis and troubleshooting tools.

  • Familiarity with Linux systems and scripting for automation (e.g., Bash, Python).

  • Strong analytical and problem-solving skills.

Preferred:

  • Experience with large-scale HPC clusters or AI/ML infrastructure.

  • Knowledge of RDMA, MPI, and low-latency networking concepts.

  • Certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent.

  • Experience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform).

Soft Skills:

  • Strong communication and collaboration skills.

  • Ability to work independently and handle complex technical challenges.

  • Detail-oriented with a proactive approach to problem-solving.

Additional Information

What does Mirantis offer you?

  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge, open-source innovation;
  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.

We are a Leader for Container Management in G2 (#2 after AWS)!

Similar Jobs

14 Minutes Ago
Remote or Hybrid
79K-135K Annually
Mid level
79K-135K Annually
Mid level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Develop and maintain project schedules and resource plans; coordinate tasks, milestones, and meetings; monitor progress, budgets, and risks; produce status reports; implement process improvements; ensure projects meet scope, timeline, and quality requirements.
Top Skills: AgileDcma 14-PointExcelMS OfficeMs ProjectSpmWaterfall
15 Minutes Ago
Remote or Hybrid
United States
Senior level
Senior level
Cloud • Information Technology • Security • Software • Cybersecurity
Serve as a quota-carrying technical trusted advisor managing the full customer lifecycle: discovery, architecture, demos/PoCs, onboarding, adoption, and expansion. Drive technical engagements, present to stakeholders, automate routine tasks with AI, and evangelize Cloudflare technologies externally.
Top Skills: APIsCloud InfrastructureCloudflare Ai GatewayCloudflare VectorizeCloudflare Workers AiDdos MitigationDnsEdge ComputingJavaScriptLlmsPythonRetrieval-Augmented Generation (Rag)RoutingWeb Security
15 Minutes Ago
Remote or Hybrid
79K-135K Annually
Senior level
79K-135K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Design, build, and maintain large-scale ETL pipelines and data warehouses (Redshift) with strong data modeling, governance, and quality controls. Optimize systems for performance and reliability, integrate and deploy AI/ML models, troubleshoot and support Redshift, collaborate with data scientists and stakeholders, and produce technical documentation and knowledge-base articles.
Top Skills: Amazon RedshiftApache AirflowApache KafkaAWSAws GlueFivetranJavaNoSQLOraclePysparkPythonPyTorchScalaScikit-LearnSparkSQLSQL ServerTensorFlow

What you need to know about the Charlotte Tech Scene

Ranked among the hottest tech cities in 2024 by CompTIA, Charlotte is quickly cementing its place as a major U.S. tech hub. Home to more than 90,000 tech workers, the city’s ecosystem is primed for continued growth, fueled by billions in annual funding from heavyweights like Microsoft and RevTech Labs, which has created thousands of fintech jobs and made the city a go-to for tech pros looking for their next big opportunity.

Key Facts About Charlotte Tech

  • Number of Tech Workers: 90,859; 6.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lowe’s, Bank of America, TIAA, Microsoft, Honeywell
  • Key Industries: Fintech, artificial intelligence, cybersecurity, cloud computing, e-commerce
  • Funding Landscape: $3.1 billion in venture capital funding in 2024 (CED)
  • Notable Investors: Microsoft, Google, Falfurrias Management Partners, RevTech Labs Foundation
  • Research Centers and Universities: University of North Carolina at Charlotte, Northeastern University, North Carolina Research Campus

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account