← Back to job listings
AN
Performance Engineer (Inference Systems)
Anthropic · New York, United States
About The Role
Join Anthropic, a leading AI safety and research company, as a Performance Engineer. In this role, you will be responsible for understanding and optimizing the entire inference system, focusing on throughput, latency, reliability, and correctness. You will work across various teams and projects, conducting cross-layer performance investigations, improving the correctness evaluation pipeline, and building observability tools. This position offers a comprehensive benefits package, including health insurance, paid parental leave, flexible time off, and competitive salary and equity packages.
- Conduct cross-layer performance investigations to identify root causes of performance gaps and quantify the value of closing them.
- Own and improve the correctness evaluation pipeline that validates model output quality across hardware platforms, and lead investigations when regressions are detected.
- Build observability, dashboards, and modeling tools that make throughput, latency, cost, reliability, correctness, and their interactions legible across the stack.
- Ability to communicate quantitative results clearly in writing to influence priorities on teams you don't manage
- Genuine interest in correctness as an engineering discipline: numerics, evaluation design, regression detection
- Proficiency in Python, with the ability to read, instrument, and contribute to large production codebases you didn’t write
- Solid data analysis skills (e.g. SQL, pandas, or similar) sufficient to turn raw telemetry into clear findings
- Hands-on performance engineering experience: profiling, roofline analysis, latency/throughput optimization, and root-cause investigation in complex production systems
- Experience with ML systems, especially training or inference infrastructure or general LLM serving stacks. Direct large-scale inference experience is a strong plus
- Familiarity with GPU/TPU/accelerator performance concepts (memory bandwidth, kernel overheads, quantization, collective communication). Reasoning about these matters more than having written kernels yourself
- Experience with reliability engineering for high-throughput services: autoscaling, load balancing, request routing, tail latency
- Experience with model evaluation or numerical regression-detection pipelines
- Experience building observability or telemetry for distributed systems
- Comfortable having impact through influence and evidence rather than direct ownership
- Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
- We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring