Skip to content
← Back to job listings

Principal Engineer (CAPE)

Crusoe Energy Systems · San Francisco, United States

External listingfull-time5 days ago

About The Role

Join Crusoe Energy, a company at the forefront of the AI revolution. As a Principal Engineer, you will work on groundbreaking technology that optimizes AI infrastructure for both speed and sustainability. Your role will involve creating a self-driving fleet of accelerators, developing a unified observability plane, and implementing closed-loop autonomy. You will have the opportunity to define the operating standard for AI at scale and work on a problem unique to Crusoe. Enjoy comprehensive health benefits, paid time off, a 401(k) match, and resources for mental wellness.

  • Concevoir et mettre en œuvre un plan d'observabilité unifiée pour corréler les signaux GPU, réseau, stockage, orchestration et charge de travail.
  • Développer et maintenir un système de gestion de flotte qui permet à des milliers d'accélérateurs de fonctionner comme un seul système programmable.
  • Travailler sur l'autonomie en boucle fermée, en diagnostiquant, décidant et en remédiant sans intervention humaine.
  • Experience applying ML/statistical methods to noisy operational telemetry (failure prediction, anomaly detection) is a strong plus
  • Prior exposure to zero-trust or policy-based multi-tenancy architectures is a plus but not required
  • Comfort operating in ambiguity and defining the architecture and standards for a system that doesn't exist yet — this is a 0→1 charter, not a maintenance role
  • Hands-on fluency with GPU/HPC infrastructure — GPU health telemetry, NVLink/InfiniBand/RoCE fabrics, thermal and power behavior at the hardware level
  • 10+ years building infrastructure-layer systems at scale — fleet management, distributed control planes, scheduler internals, or hardware lifecycle automation. This is a systems-builder role, not a consumer of managed cloud services
  • Strong software engineering fundamentals in at least one systems language (Go, Rust, C++, or similar) and the judgment to know when to build vs. adopt existing tooling
  • Track record of designing and shipping large-scale observability or telemetry platforms that correlate signals across compute, network, and storage layers
  • Deep experience with distributed systems design: consensus, state reconciliation, closed-loop automation, and systems that make autonomous decisions against live production infrastructure

This is an external listing. JobSpring does not represent or verify the employer. Report this listing