← Back to job listings
SO
Senior AI Validation Engineering Manager
Sonatus · Sunnyvale, United States
About The Role
Join Sonatus, a leading company in the automotive technology sector, as a Senior AI Validation Engineering Manager. In this role, you will build and lead the AI Validation function, responsible for testing, evaluating, and governing AI models and capabilities in our software-defined vehicle and cloud platforms. You will define and drive Sonatus's AI validation strategy, lead a high-performing AI validation team, and collaborate with engineering, product, and safety stakeholders. This is an exciting opportunity to make AI quality and safety a competitive advantage in the automotive industry.
- Definir y dirigir la estrategia de validación de IA de Sonatus, identificando brechas en el desarrollo, prueba, implementación y gobernanza de modelos.
- Liderar, contratar, mentorizar y hacer crecer una organización de validación de IA de alto rendimiento, estableciendo procesos de ingeniería escalables y dirección técnica.
- Ser el responsable de la estrategia de validación de extremo a extremo para modelos de IA/ML, LLMs, pipelines RAG y flujos de trabajo de IA agentiva.
- Practical experience with RAG evaluation frameworks such as RAGAS, including evaluation tuning, embedding optimization, retrieval quality improvement, and production-scale LLM evaluation pipelines
- Strong experience testing cloud-native platforms and cloud-managed embedded products, including end-to-end system validation
- Hands-on experience developing, deploying, testing, or operating AI/ML systems, with strong expertise in modern ML workflows, neural networks, and MLOps
- 10+ years of experience in software or systems engineering—including embedded, cloud, networking, security, or automotive domains—with 3+ years leading high-performing engineering or QA organizations
- Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field required (MS preferred)
- Strong expertise with hybrid evaluation methodologies combining deterministic validation (citation grounding, structural validation, exact matching) and probabilistic LLM-as-a-Judge techniques (faithfulness, answer relevance, context precision, task completion)
- Proven experience designing and implementing scalable evaluation frameworks for AI systems, including multi-step agentic workflows, regression testing, benchmarking, and automated quality scoring
- Experience designing systems that verify external knowledge claims and ensure responses are grounded in traceable citations and trusted data sources
- Experience establishing AI governance, safety, compliance, and Responsible AI practices for enterprise or safety-critical systems
- Experience validating hallucination, grounding, citation accuracy, bias, fairness, toxicity, and factual consistency in production LLM applications
- Proficiency in Python, Linux, shell scripting, modern test frameworks (PyTest, Playwright, Behave), and engineering productivity tools such as Jenkins and JIRA
- Deep understanding of LLMs, RAG architectures, vector databases, embeddings, retrieval optimization, and agentic AI frameworks such as LangGraph or equivalent orchestration platforms
- Experience validating AI or agentic systems in safety-critical or regulated industries (automotive, aerospace, medical)
- Track record building and scaling an AI test/evaluation platform or developer experience used by multiple teams (frameworks, reusable components, reference implementations)
- Demonstrated wins moving AI testing practices from ad hoc to standardized, organization-wide adoption, with measurable impact on cycle time, quality, or reliability
- Experience implementing enterprise-grade AI governance (auditability, monitoring, policy enforcement) in production systems
- Deep experience evaluating LLM and RAG systems at scale, including agentic workflows, RAGAS-based evaluation, citation verification, hallucination detection, groundedness, faithfulness, answer relevance, tool/task correctness, and automated regression testing across offline and online feedback loops
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring