Skip to content
← Back to job listings

Research Engineer (Model Evaluations)

Anthropic · New York, United States

External listingfull-time25 days ago

About The Role

Join Anthropic, a leading AI safety and research company, as a Research Engineer focused on model evaluations. In this role, you will turn abstract concepts of intelligence into measurable metrics, design and implement evaluations of Claude's capabilities, and build the infrastructure to run them at scale. You will collaborate closely with researchers throughout the lifecycle of new capabilities and communicate results to internal and external stakeholders.

  • Design and implement evaluations across the full spectrum of Claude's capabilities and personality, and build the infrastructure that runs them reliably at scale.
  • Partner closely with researchers throughout the lifecycle of a new capability — from defining what to measure, to running the eval against live training checkpoints, to interpreting the results.
  • Run experiments to characterize how prompting, sampling, and scaffolding choices affect results on internal and industry benchmarks.
  • Care about the societal impacts of your work and an interest in steering powerful AI to be safe and beneficial
  • Experience building or operating distributed systems, data pipelines, or other infrastructure that needs to be reliable at scale
  • Strong Python programming skills, including production or research infrastructure
  • Comfort operating in an on-call or production-support capacity when training runs are live
  • Clear written and verbal communication, especially when explaining technical results to non-specialists
  • Hands-on experience using large language models such as Claude, including prompting, sampling, and scaffolding
  • Background in data visualization and a track record of building dashboards people actually trust and use
  • Experience developing robust evaluation metrics for language models
  • Experience with observability, monitoring, or experiment-tracking systems
  • Background in statistics and experimental design
  • Experience with large-scale dataset sourcing, curation, and processing
  • Experience running or supporting ML training infrastructure
  • A bias toward picking up slack and operating flexibly across team boundaries
  • Enjoy pair programming — we love to pair
  • Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
  • Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
  • Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices
  • We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed
  • Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work

This is an external listing. JobSpring does not represent or verify the employer. Report this listing