Skip to content
← Back to job listings

Staff+ Software Engineer (Inference Runtime)

Anthropic · New York, United States

External listingfull-timeabout 2 months ago

About The Role

all qualifications are necessary to apply for this position.

  • Set technical direction for the team's architecture, roadmap, and shared runtime of the inference serving stack.
  • Own and evolve the accelerator-agnostic runtime, including hands-on work in a performance-sensitive Rust and Python codebase.
  • Drive efficient accelerator usage across GPU, TPU, and Trainium, and build the runtime's validation surface around partitioned builds.
  • This role is for someone who has been the technical anchor of a platform with many internal consumers, who thinks in systems and feedback loops, and who gets real satisfaction from building abstractions that hold up as the system scales another order of magnitude
  • Deep background in systems engineering or ML infrastructure, with the ability to go hands-on with performance profiling, latency and throughput optimization, and systems debugging at scale
  • Strong written and verbal communication, and the ability to influence technical direction without formal authority
  • A track record of defining and using engineering metrics to drive improvement: you've set SLOs on platform surfaces, and driven escape rates, release times, latency, or throughput in a measurable direction
  • Have significant software engineering experience, with a strong background in high-performance, large-scale distributed systems serving millions of users
  • Real depth in at least one accelerator ecosystem (CUDA/GPU, TPU, or Trainium/AWS Neuron) and genuine appetite to keep the runtime agnostic across all of them
  • Experience driving technical alignment across organizational boundaries, advocating for your team's needs while contributing to shared infrastructure
  • Prior tech lead experience on a developer productivity or platform engineering team at a fast-growing AI/ML company
  • Experience with deterministic or simulation-based testing for hardware-dependent systems
  • Experience with CI/CD systems at scale, particularly for workloads involving accelerator hardware
  • Familiarity with Kubernetes-based development and job scheduling environments
  • Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
  • Background operating production as a validation surface at scale: shadow traffic, canary populations, automated baseline comparison, fast rollback
  • Experience with ML compiler toolchains (XLA, Triton, NeuronX) or accelerator driver/firmware management at scale
  • 8+ years of software engineering experience, with significant time as the technical lead or anchor on a platform, inference runtime, or ML infrastructure team
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
  • We encourage you to apply even if you do not believe you meet every single qualification. Not

This is an external listing. JobSpring does not represent or verify the employer. Report this listing