← Back to job listings
TA
Staff Software Engineer (Inference / Compute Infrastructure Engineering)
Together AI · San Francisco, United States
About The Role
Join our team as a Staff Software Engineer focused on Inference and Compute Infrastructure Engineering. In this role, you will build systems that treat infrastructure as software, owning the software state machines that provision hardware, bring it into service, and manage its full lifecycle. You will design and implement the software that models the full lifecycle of a physical host, create a self-service API for the inference team, automate self-healing, and ensure the reliability of the pipeline. You will also partner with the inference/ML platform team and engineer the platform like software.
- Concevoir et mettre en œuvre le logiciel qui modélise l'ensemble du cycle de vie d'un hôte physique, de la découverte à la mise hors service.
- Développer des API déclaratives et un plan de contrôle afin que l'équipe d'inférence puisse demander, mettre à l'échelle et détruire des clusters d'inférence par un seul appel API.
- Automatiser l'auto-réparation : détecter les nœuds dégradés ou défaillants, les drainer en toute sécurité, déclencher la réparation ou le remplacement.
- Experience with event-driven systems — designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts
- Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines)
- Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution
- A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship
- Strong software engineering background in Go, Python, Rust, or similar — you write and test real software for a living
- Exposure to bare-metal provisioning (PXE/iPXE, Redfish/IPMI, BMC) and/or networking fundamentals (VLANs, BGP, fabric design), or GPU/accelerator infrastructure
- Experience with GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE)
- Prior work at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization
- Systems programming in Rust or Go
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring