Skip to content
← Back to job listings

Senior HPC Networking Engineer

Mirantis · Remote, United States

IT - Network / Systems / DB AdminSenior LevelRemoteExternal listingfull-timeabout 7 hours ago

About The Role

Employment Type: Full-time

Role Overview

We are seeking a highly skilled Senior HPC Networking Engineer to design, deploy, manage, and troubleshoot high-performance networking environments. The ideal candidate will have deep expertise in InfiniBand technologies, strong general networking knowledge, and hands-on experience with Fortinet solutions. You will play a critical role in ensuring the performance, reliability, and scalability of HPC infrastructure. 

Key Responsibilities

  • Design, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics.
  • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
  • Manage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations.
  • Perform performance tuning, monitoring, and capacity planning for HPC networking systems.
  • Implement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).
  • Diagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments.
  • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.
  • Develop and maintain documentation for network architecture, configurations, and operational procedures.
  • Participate in on-call rotations and provide escalation support for critical incidents.
  • Lead or contribute to network upgrades, migrations, and new deployments.
  • You will actively troubleshoot and resolve daily customer incident tickets to keep massive GPU training runs moving.

Build, operate, and scale next-generation GPU infrastructure

  • You will be hands-on on the front lines driving daily triage, incident resolution, and SLA management for the world’s most advanced NVIDIA clusters, InfiniBand/RoCE fabrics, and AI workloads.

Build the playbook, then grow into the platform

  • Designed for engineers energized by standing up new operations from the ground up, this role offers a direct trajectory from operationalizing bare-metal clusters to driving platform and AI capabilities as we scale.

Required

  • 5+ years of experience in network engineering, with a focus on HPC or data center environments.
  • Strong hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA).
  • Solid understanding of networking fundamentals: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.
  • Proven experience deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).
  • Experience with network performance analysis and troubleshooting tools.
  • Familiarity with Linux systems and scripting for automation (e.g., Bash, Python).
  • Strong analytical and problem-solving skills.

Preferred

  • Experience with large-scale HPC clusters or AI/ML infrastructure.
  • Knowledge of RDMA, MPI, and low-latency networking concepts.
  • Certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent.
  • Experience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform).

Soft Skills

  • Strong communication and collaboration skills.
  • Ability to work independently and handle complex technical challenges.
  • Detail-oriented with a proactive approach to problem-solving.

What We Offer

  • Opportunity to work on cutting-edge HPC infrastructure.
  • Collaborative and innovative work environment.
  • Competitive salary and benefits package.
  •  
  • What does Mirantis offer you?
  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge, open-source innovation;
  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.

We are a  Leader for Container Management  in G2 (#2 after AWS)!

This is an external listing. JobSpring does not represent or verify the employer. Report this listing