← Back to job listings
SA
Senior Machine Learning Engineer (Agent Oversight)
Scale AI · New York, United States
About The Role
Join Scale AI as a Senior Machine Learning Engineer focused on Agent Oversight. In this role, you will drive the end-to-end lifecycle of our production agents, ensuring their reliability and continuous improvement. You will build observability tools, design evaluation frameworks, and develop improvement loops. You will collaborate closely with product managers, customers, and other engineering teams to translate requirements into robust platform capabilities. This position offers comprehensive health coverage, personal and career growth opportunities, and parental support.
- Contribuer à la construction d'outils d'observabilité pour surveiller le comportement des agents en production.
- Concevoir des méthodologies d'évaluation et des métriques pour les applications agentiques, et travailler avec la plateforme pour les automatiser à grande échelle.
- Construire, expédier et posséder des systèmes d'apprentissage automatique qui détectent les dérives, les anomalies ou les désalignements dans le comportement des agents en production.
- Rigorous approach to experimentation: clear hypotheses, real statistical grounding, and results that hold up under scrutiny
- Comfortable partnering with software engineers to productionize research and experimental work, not just deliver a one-off analysis
- Hands-on experience with LLMs and agent architectures — tool use, planning, multi-agent orchestration
- 5+ years of experience as an ML engineer or applied scientist, ideally on a production ML or LLM-powered system — not just consuming a third-party ML API within a feature
- Developing new methods, reward models, or model training/fine-tuning approaches
- Strong grounding in at least two of the following:
- Track record of collaborating across functions (Product, Forward Deployed Engineering, etc.) to navigate ambiguous requirements and bring them to production
- Design experience for agent systems (architecture, orchestration, tool use)
- Building or scaling evaluation, monitoring, or continuous-learning infrastructure for ML/agentic systems
- Gives direct, substantive feedback on designs and code, and takes it the same way — and mentors others as they grow
- Experience building or contributing to RLHF, SFT, or other fine-tuning/RL workflows, reward modeling, or verifiable-reward systems
- Experience with model or systems optimization (e.g., latency, cost, or inference efficiency)
- Published research, open-source contributions, or patents in agentic systems, LLMs, or applied ML
- Track record of taking a novel method from prototype to something running reliably in production, navigating ambiguity along the way
- Experience working in regulated or enterprise contexts
- Experience reviewing others’ technical designs or mentoring engineers at a senior/staff level
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring