← Back to job listings
CE
Principal Engineer (CAPE)
Crusoe Energy Systems · San Francisco, United States
About The Role
Join Crusoe Energy, a company at the forefront of the AI revolution. As a Principal Engineer, you will work on groundbreaking technology that optimizes AI infrastructure for both speed and sustainability. Your role will involve creating a self-driving fleet of accelerators, developing a unified observability plane, and implementing closed-loop autonomy. You will have the opportunity to define the operating standard for AI at scale and work on a problem unique to Crusoe. Enjoy comprehensive health benefits, paid time off, a 401(k) match, and resources for mental wellness.
- Concevoir et mettre en œuvre un plan d'observabilité unifiée pour corréler les signaux GPU, réseau, stockage, orchestration et charge de travail.
- Développer et maintenir un système de gestion de flotte qui permet à des milliers d'accélérateurs de fonctionner comme un seul système programmable.
- Travailler sur l'autonomie en boucle fermée, en diagnostiquant, décidant et en remédiant sans intervention humaine.
- Experience applying ML/statistical methods to noisy operational telemetry (failure prediction, anomaly detection) is a strong plus
- Prior exposure to zero-trust or policy-based multi-tenancy architectures is a plus but not required
- Comfort operating in ambiguity and defining the architecture and standards for a system that doesn't exist yet — this is a 0→1 charter, not a maintenance role
- Hands-on fluency with GPU/HPC infrastructure — GPU health telemetry, NVLink/InfiniBand/RoCE fabrics, thermal and power behavior at the hardware level
- 10+ years building infrastructure-layer systems at scale — fleet management, distributed control planes, scheduler internals, or hardware lifecycle automation. This is a systems-builder role, not a consumer of managed cloud services
- Strong software engineering fundamentals in at least one systems language (Go, Rust, C++, or similar) and the judgment to know when to build vs. adopt existing tooling
- Track record of designing and shipping large-scale observability or telemetry platforms that correlate signals across compute, network, and storage layers
- Deep experience with distributed systems design: consensus, state reconciliation, closed-loop automation, and systems that make autonomous decisions against live production infrastructure
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring