Skip to content
← Back to job listings

AI Evaluation Engineer

Zafin · Toronto, Canada

External listingfull-time5 days ago

About The Role

Join Zafin, a leading provider of banking software solutions. As an AI Evaluation Engineer, you will ensure the accuracy, reliability, safety, and production-readiness of AI agent solutions within regulated banking environments. You will design and execute evaluation scenarios, validate AI agent behavior, conduct regression evaluation, identify quality issues, and support production-readiness decisions. You will work closely with various teams to ensure evaluation reflects intended business logic, regulatory requirements, technical standards, and real-world operating conditions.

  • Design and execute structured evaluation scenarios that validate AI agent accuracy, reliability, safety, compliance, and business outcomes.
  • Identify, document, prioritize, and track quality issues, defects, and production risks, and support production-readiness decisions through objective evaluation evidence.
  • Collaborate with cross-functional teams to ensure evaluation reflects intended business logic, regulatory requirements, technical standards, and real-world operating conditions.
  • Level and scope of responsibility will be determined based on demonstrated technical capability, evaluation expertise, leadership experience, independence, and ability to influence quality outcomes
  • Understanding of privacy, security, governance, and regulatory considerations relevant to enterprise AI
  • Working knowledge of SDLC, CI/CD, automated evaluation, AI observability, and engineering delivery practices
  • Experience leading evaluation activities, mentoring technical professionals, or coordinating quality initiatives is advantageous for more senior levels
  • Proficiency with Python, SQL, or similar tools supporting evaluation and analysis
  • Typically 3–10+ years of relevant experience in software quality engineering, AI evaluation, AI quality engineering, machine learning evaluation, software testing, or related disciplines
  • Familiarity with benchmark management, evaluation tooling, quality automation, and AI engineering workflows
  • Degree in Computer Science, Software Engineering, Data Science, Artificial Intelligence, or related discipline, or equivalent practical experience
  • Experience with structured software testing, regression evaluation, production-readiness assessment, and quality engineering
  • Strong understanding of AI evaluation, large language model behaviour, reasoning quality, hallucination detection, safety, instruction adherence, factual accuracy, and business correctness
  • Experience evaluating LLMs, RAG systems, AI agents, or agentic AI platforms
  • Experience with AI evaluation platforms such as LangSmith, OpenAI Evals, or comparable tools
  • Experience integrating automated evaluation into CI/CD or MLOps workflows
  • Experience with model observability, behavioural-drift detection, or AI production monitoring
  • Banking, financial services, or other regulated industry experience
  • Experience leading technical teams, quality initiatives, or engineering improvement programs

This is an external listing. JobSpring does not represent or verify the employer. Report this listing