← Back to job listings
AN
Staff + Senior Software Engineer (Inference Deployment)
Anthropic · New York, United States
About The Role
Join Anthropic, a leading AI safety and research company, as a Staff + Senior Software Engineer in Inference Deployment. In this role, you will design and build the deployment infrastructure that moves inference code from merge to production. You will work on optimizing deployment scheduling, improving observability, and driving down cycle time from code merge to production. You will also partner with teams across the Inference organization to integrate deployment automation with their systems.
- Concevoir et construire l'infrastructure de déploiement qui déplace le code d'inférence de la fusion à la production.
- Améliorer la planification des déploiements en tenant compte de la capacité pour maximiser le débit de déploiement.
- Évoluer l'intégration de l'automatisation du déploiement avec les systèmes de validation, d'autoscaling et de routage des modèles.
- Comfort working across the stack — from backend services and databases to CLI tools and web UIs
- Strong communication skills and the ability to work closely with oncall engineers, model teams, and infrastructure partners
- Experience building deployment, release, or delivery infrastructure where resource constraints (fleet capacity, network bandwidth, hardware availability, coordinated rollout windows) shape the design
- A track record of building automation that measurably improves deployment velocity and reliability
- Proficiency with Kubernetes-based deployments, rolling update mechanics, and container orchestration
- Strong software engineering skills, including experience designing systems that manage complex state machines and multi-stage pipelines
- Experience with progressive delivery in systems with long validation cycles: canary/soak testing, blue-green deployments, traffic shifting, automated rollback
- Bachelor’s degree or an equivalent combination of education, training, and/or experience
- Experience at companies with large-scale release engineering challenges (mobile release trains, monorepo deployments, multi-datacenter rollouts)
- Experience with ML inference or training infrastructure deployment, particularly across multiple accelerator types (GPU, TPU, Trainium)
- 5+ years of experience building deployment, release, or delivery infrastructure at scale
- Experience with Python and/or Rust in production systems
- Background in capacity planning or resource-constrained scheduling (e.g., bin-packing, fleet management, job scheduling with hardware affinity)
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring