Skip to content
← Back to job listings

Staff Infrastructure Engineer (Cluster Infrastructure)

Anthropic · London, United Kingdom

External listingfull-time23 days ago

About The Role

Join Anthropic, a leading AI safety and research company. As a Staff Infrastructure Engineer, you will be responsible for the technical strategy and roadmap for agent-driven cluster lifecycle management. You will collaborate with various teams to ensure timely ingestion of new compute capacity, align with partner teams on physical build-out, and define strategies for cluster scalability and fault tolerance. You will also establish operational-excellence practices and support the growth of engineers through mentorship and coaching.

  • Own the technical strategy and roadmap for agent-driven cluster lifecycle management, including provisioning, updates, and decommissioning.
  • Collaborate with security owners to ensure clusters are provisioned secure-by-default and define and drive strategy on cluster scalability, homogeneity, and fault tolerance.
  • Establish and evolve operational-excellence practices, including incident response, postmortem culture, and on-call health, while supporting the growth of engineers through technical mentorship and coaching.
  • 10+ years of software engineering experience, including time as a technical lead setting direction for a team
  • Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure)
  • Ability to build alignment across senior stakeholders and communicate effectively at all levels
  • Strong proficiency in at least one systems language (e.g., Rust, Go, or Python), IaC proficiency with Terraform
  • Track record of leading complex, multi-quarter technical initiatives spanning multiple teams or systems
  • Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
  • Experience with cluster security: pod security standards and admission control, RBAC and least-privilege IAM, node and container hardening, supply-chain/image provenance
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
  • Deep experience with infrastructure-as-code (Terraform, Atlantis), workflow orchestration (Temporal, Argo Workflows)
  • Experience with cluster and host networking: CNI (Cilium), eBPF, NetworkPolicy, multi-NIC, sFlow, service mesh (Istio/Envoy/Linkerd, mTLS)
  • Experience with cloud networking: VPC design and peering, Shared VPC/Transit Gateway, Cloud Interconnect/Direct Connect, Cloud NAT, cross-cloud private connectivity, BGP and route control, edge load balancing and DDoS mitigation (Cloud Armor / AWS Shield)
  • Skill in quickly understanding systems design tradeoffs and keeping track of rapidly evolving software systems
  • Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
  • Depth in one or more of: Kubernetes internals, cluster provisioning and management systems, cluster orchestration systems (Mesos, Borg-like)
  • Experience operating large-scale compute infrastructure at hyperscale (100+ clusters, 10K+ nodes)
  • We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed

This is an external listing. JobSpring does not represent or verify the employer. Report this listing