Skip to content
← Back to job listings

Staff+ Software Engineer (Safeguards ML Infrastructure)

Anthropic · San Francisco, United States

External listingfull-time24 days ago

About The Role

Join Anthropic, an AI safety and research company, as a Staff+ Software Engineer in the Safeguards ML Infrastructure team. You will design, build, and operate the production infrastructure that powers Claude's safety systems. Your responsibilities will include owning and operating the production serving infrastructure, defining and maintaining SLOs, building observability systems, and leading incident response. You will also participate in on-call and operational-duty rotations, build automation and tooling, and work closely with ML researchers to productionize new safety techniques.

  • Design, build, and deploy backend services that are critical safety pieces on the token sampling and generation path.
  • Own and operate the production serving infrastructure for those services across multiple deployment platforms (1P, AWS Bedrock, GCP Vertex).
  • Define and maintain SLOs, build observability and alerting systems, and lead incident response for infrastructure on the critical path of every Claude request.
  • Have hands-on experience deploying and operating on cloud platforms (AWS, GCP) at scale
  • Have a strong foundation in distributed systems: replication, consistency tradeoffs, failure modes, and SLO management under load
  • Have meaningful on-call experience for production systems, including incident response and postmortem-driven improvements
  • Have designed, built, and operated high QPS systems at global scale
  • Approach infrastructure as a platform -- building systems and abstractions that other engineers build on, rather than point solutions for a single team's needs
  • Are proficient in Python; experience with Rust is a plus but not required
  • Have a desire to close the gap where nobody has yet raised their hand, even if it requires manually hand holding processes until automation and tooling can be built
  • Familiarity with LLM inference systems and the operational characteristics of transformer-based models
  • A demonstrated history of reducing operational toil through automation, including transitioning teams from manual deployment processes to self-serve pipelines
  • We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed
  • Experience building deployment and rollout systems with canary analysis, automated validation, or progressive rollout controls
  • Have 8+ years of industry software engineering experience
  • Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices

This is an external listing. JobSpring does not represent or verify the employer. Report this listing