← Back to job listings
AN
Staff+ Software Engineer (Safeguards Evals)
Anthropic · New York, United States
About The Role
Join Anthropic, a leading AI safety and research company, as a Staff Software Engineer focused on Safeguards Evaluations. In this role, you will build the evaluation infrastructure that measures the performance of our investigative agent across various harm areas. Your work will directly impact the trust we place in our automated abuse detection systems and guide our investment in improving them. You will have the opportunity to work at the intersection of applied ML research and engineering, and your contributions will be critical to ensuring the safety and reliability of our AI systems.
- Construire et posséder l'infrastructure d'évaluation pour un système d'investigation agentique, en définissant des métriques, des cas de test et des approches de notation.
- Construire des ensembles de données d'évaluation de haute qualité représentant des abus réels dans divers domaines de préjudice, en s'appuyant sur des modèles de trafic réels et une génération synthétique.
- Mesurer la performance de l'agent de bout en bout (précision/rappel de détection, qualité d'investigation, robustesse) et conduire l'amélioration continue dans les domaines de préjudice les plus difficiles.
- Ability to move fluidly between research prototyping and production-quality code
- Experience working with LLMs and a working understanding of their capabilities and failure modes — especially agentic systems with tool use and multi-step reasoning
- Strong data analysis skills — you can draw reliable insights from large datasets
- Experience building and maintaining data pipelines
- Ability to translate ambiguous problems into concrete, testable experiments
- Proficiency in Python and comfort working across the stack
- Extensive experience in trust and safety, content moderation, or abuse detection systems
- Experience with distributed systems or large-scale data processing
- We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed
- Experience in red teaming, adversarial testing, or jailbreak research on AI systems
- Expertise in building or contributing to agent evaluation frameworks, benchmarks, or automated grading systems
- Experience with prompt engineering or building LLM-powered applications
- Experience with synthetic data generation or data augmentation
- 8+ years of industry software engineering experience
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring