← Back to job listings
ZA
AI Evaluation Engineer
Zafin · Toronto, Canada
About The Role
Join Zafin, a leading provider of banking software solutions. As an AI Evaluation Engineer, you will ensure the accuracy, reliability, safety, and production-readiness of AI agent solutions within regulated banking environments. You will design and execute evaluation scenarios, validate AI agent behavior, conduct regression evaluation, identify quality issues, and support production-readiness decisions. You will work closely with various teams to ensure evaluation reflects intended business logic, regulatory requirements, technical standards, and real-world operating conditions.
- Design and execute structured evaluation scenarios that validate AI agent accuracy, reliability, safety, compliance, and business outcomes.
- Identify, document, prioritize, and track quality issues, defects, and production risks, and support production-readiness decisions through objective evaluation evidence.
- Collaborate with cross-functional teams to ensure evaluation reflects intended business logic, regulatory requirements, technical standards, and real-world operating conditions.
- Level and scope of responsibility will be determined based on demonstrated technical capability, evaluation expertise, leadership experience, independence, and ability to influence quality outcomes
- Understanding of privacy, security, governance, and regulatory considerations relevant to enterprise AI
- Working knowledge of SDLC, CI/CD, automated evaluation, AI observability, and engineering delivery practices
- Experience leading evaluation activities, mentoring technical professionals, or coordinating quality initiatives is advantageous for more senior levels
- Proficiency with Python, SQL, or similar tools supporting evaluation and analysis
- Typically 3–10+ years of relevant experience in software quality engineering, AI evaluation, AI quality engineering, machine learning evaluation, software testing, or related disciplines
- Familiarity with benchmark management, evaluation tooling, quality automation, and AI engineering workflows
- Degree in Computer Science, Software Engineering, Data Science, Artificial Intelligence, or related discipline, or equivalent practical experience
- Experience with structured software testing, regression evaluation, production-readiness assessment, and quality engineering
- Strong understanding of AI evaluation, large language model behaviour, reasoning quality, hallucination detection, safety, instruction adherence, factual accuracy, and business correctness
- Experience evaluating LLMs, RAG systems, AI agents, or agentic AI platforms
- Experience with AI evaluation platforms such as LangSmith, OpenAI Evals, or comparable tools
- Experience integrating automated evaluation into CI/CD or MLOps workflows
- Experience with model observability, behavioural-drift detection, or AI production monitoring
- Banking, financial services, or other regulated industry experience
- Experience leading technical teams, quality initiatives, or engineering improvement programs
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring