← Back to job listings
AN
Senior Staff Software Engineer (Node Infrastructure)
Anthropic · London, United Kingdom
About The Role
Join Anthropic, a leading AI safety and research company. As a Senior Staff Software Engineer in the Node Infrastructure team, you will own the technical strategy and roadmap for node lifecycle management, drive cross-team initiatives, and design and operate systems for hardware health and repair. You will work closely with cloud providers and internal teams to shape long-term infrastructure strategy and support the growth of engineers through mentorship and coaching. Enjoy comprehensive benefits, a competitive salary, and the opportunity to make a significant impact in the field of AI.
- Own the technical strategy and roadmap for node lifecycle management, including ingestion, bring-up, health checking, and automated repair.
- Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families, ensuring the hardest problems get solved.
- Establish and evolve operational excellence practices, including incident response, postmortem culture, and on-call, while supporting the growth of engineers through technical mentorship and coaching.
- Hands-on experience with machine learning accelerators (GPUs, TPUs, or Trainium)
- Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure)
- Track record of leading complex, multi-quarter technical initiatives that span multiple teams or systems
- Strong proficiency in at least one systems language (e.g., Rust, Go, or Python), IaC proficiency with Terraform
- Ability to build alignment across senior stakeholders and communicate effectively at all levels
- Experience managing large scale compute infrastructure at hyperscale (10K+ nodes), including capacity management and efficiency
- Depth in one or more of: Kubernetes internals (scheduler, autoscaler, kubelet, Karpenter), cluster orchestration systems (Mesos, Borg-like), or node provisioning pipelines
- Low-level systems experience: kernel, virtualization, device drivers, firmware, or hardware health/diagnostics daemons
- Familiarity with high-performance networking (EFA, RDMA, InfiniBand) for distributed ML workloads
- Contributions to relevant open-source projects (Kubernetes, Linux kernel, container runtimes, etc.)
- Demonstrated ownership of production reliability for high-throughput, latency-sensitive systems
- Skill in quickly understanding systems design tradeoffs and keeping track of rapidly evolving software systems
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
- Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
- We encourage you to apply even if you do not believe you meet every single qualification
- 12+ years of software engineering experience, including time as a technical lead setting direction for a team
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring