← Back to job listings
CE
Staff Software Engineer (Cloud Infrastructure)
Crusoe Energy Systems · San Francisco, United States
About The Role
Join Crusoe's Fleet Operations team as a Staff Software Engineer focused on cloud infrastructure. In this role, you will be responsible for the advanced diagnosis, maintenance, and repair of high-performance GPU compute clusters. You will work closely with data center operations, engineering, and vendors to support cutting-edge infrastructure featuring the latest NVIDIA and AMD GPUs. This position is critical for maintaining the health and scalability of Crusoe's rapidly growing GPU fleet.
- Automate deep-level diagnosis and troubleshooting of hardware faults within GPU racks and high-density compute systems.
- Develop software to troubleshoot and support GPU platforms including NVIDIA A100, H200, GB200, B200, B300 and AMD 350X / 355X.
- Implement and execute preventative maintenance procedures to improve fleet reliability and extend hardware lifespan.
- Experience working with enterprise server hardware, power delivery, and cooling systems
- Strong Linux experience (Ubuntu, Rocky Linux, CentOS) using the command line for diagnostics and testing
- Experience supporting NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X series platforms
- Ability to work independently in a fast-paced data center or operations environment
- Strong analytical and problem-solving skills
- Deep understanding of GPU architectures and hands-on experience with GPU-based systems
- Proficiency with GPU and system diagnostic tools such as NVIDIA DCGM and NVIDIA field diagnostic utilities
- Familiarity with high-speed interconnects such as InfiniBand, NVLink, and RDMA over Converged Ethernet (RoCE)
- Excellent communication and collaboration skills
- Proven experience diagnosing and repairing high-density, rack-mounted compute hardware in production environments
- Ability to code in Golang
- Technical certification or Associate’s/Bachelor’s degree in Electrical Engineering, Computer Science, or a related field or demonstrated experience
- Experience working directly with hardware vendors and escalations
- Background in large-scale GPU fleet operations or hyperscale data center environments
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring