← Back to job listings
AN
Senior Staff+ Research Engineer (Safeguards Labs)
Anthropic · New York, United States
About The Role
Join Anthropic, a leading AI safety and research company, as a Senior Staff+ Research Engineer. In this role, you will define and execute the Labs research agenda, lead and contribute to research projects focused on detecting misuse of AI models, and develop prototypes for real-time safeguards. You will have substantial latitude over your work and high leverage on the team's direction. The position offers a comprehensive benefits package, including health insurance, paid parental leave, flexible paid time off, and competitive salary and equity packages.
- Definir y ejecutar la agenda de investigación del laboratorio, incluyendo la planificación y ejecución de proyectos de investigación.
- Liderar y contribuir a proyectos de investigación que investiguen nuevos métodos para detectar el uso indebido de modelos de lenguaje, identificando organizaciones y cuentas maliciosas.
- Desarrollar y iterar prototipos que podrían eventualmente alimentar señales en el camino de salvaguardias en tiempo real, colaborando con ingenieros en la transferencia de tecnología.
- Have working familiarity with how large language models operate — sampling, prompting, training — even if LLMs aren't your primary background
- Care about the societal impacts of AI and want your work to directly reduce real-world harm
- Are proficient in Python and comfortable working with large datasets
- Have a track record of independently driving research projects from ambiguous problem statements to concrete results, ideally in AI, ML, security, integrity, or a related technical field
- Are comfortable scoping your own work and switching between research, engineering, and analysis as a project demands
- Experience with red teaming, jailbreak research, or interpretability methods like steering vectors
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
- A history of taking research prototypes and transferring them into production systems
- Experience building and training machine learning models, including classifiers for abuse, fraud, integrity, or security applications
- Experience with agentic environments and evaluating model behavior in them
- Knowledge of evaluation methodologies for language models and experience designing evals
- Background in trust and safety, integrity, fraud detection, threat intelligence, or adversarial ML
- Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
- We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed
- Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring