Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

AI Evaluation Engineer

Zafin
CompanyZafin
CategoryEngineering
LocationToronto
RemoteOn-site (inferred)
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted23 Jul 2026
Last verified30 Jul 2026
SourceEmployer career page (greenhouse)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Zafin is an AI platform company helping regulated institutions modernize how critical work is designed, governed, and delivered. Our technology enables organizations to move faster while maintaining the governance, accountability, and control required in highly regulated environments. Our portfolio includes Zafin AIOS , an agent orchestration platform for governed AI work; the Zafin Banking Platform , which helps banks modernize product, pricing, offers, billing, loyalty, and relationship management; and Zafin IO , an integration platform that connects data, systems, and workflows across complex enterprise environments. Headquartered in Toronto, Canada, Zafin partners with leading financial institutions across North America, Europe, the Middle East, Africa, and Asia-Pacific. As AI transforms the future of financial services, we're building the platforms that help regulated organizations adopt AI responsibly and at scale. What’s the Opportunity?   The AI Evaluation Engineer ensures AI agent solutions are accurate, reliable, safe, and production-ready within regulated banking environments. The role provides objective, evidence-based evaluation of AI agent behaviour against defined business, quality, risk, performance, and regulatory criteria, helping ensure AI solutions deliver consistent, trusted outcomes in production. Working within the Reliability Testing phase of the AIOS (Zafin's AI Operating System) delivery lifecycle, the role designs, executes, and leads evaluation activities that validate AI agent behaviour across the full development lifecycle. This includes developing meaningful evaluation scenarios, identifying defects and failure modes, monitoring quality across releases, and providing actionable feedback that continuously improves AI agent reliability. The role helps operationalize AIOS's principle of reliability first, velocity second through disciplined evaluation and objective production-readiness decisions. The AI Evaluation Engineer works closely with Agent Engineers, Industry Consultants, Agent Architects, AI Knowledge & Governance, Product, and Delivery teams to ensure evaluation reflects intended business logic, regulatory requirements, technical standards, and real-world operating conditions. Depending on experience and level, the role may also lead evaluation activities, coach other evaluation engineers, improve evaluation practices, and drive quality improvements across multiple AI agent initiatives. Ultimately, the role helps ensure AI agent capabilities earn and maintain customer trust by delivering consistent, reliable, and explainable outcomes in production. What Will You Do?   Design and execute structured evaluation scenarios that validate AI agent accuracy, reliability, safety, compliance, and business outcomes. Validate AI agent behaviour against approved business rules, policies, technical requirements, source knowledge, and expected outcomes. Conduct regression evaluation across releases and monitor behavioural drift, performance degradation, and newly introduced failure modes. Identify, document, prioritize, and track quality issues, defects, and production risks. Support production-readiness decisions through objective evaluation evidence and recommendations. Analyze evaluation results to identify root causes, recurring quality trends, and opportunities to improve prompts, workflows, knowledge, integrations, and engineering practices. Maintain reusable evaluation scenarios, benchmark datasets, expected outcomes, regression suites, and supporting evidence. Contribute to continuous improvement of evaluation methodologies, automation, tooling, and engineering feedback loops. Support investigation of production issues and validate corrective actions. Partner with Agent Engineering, Industry Consultants, Agent Architects, AI Knowledge & Governance, Product, and Delivery teams to ensure evaluation reflects business requirements and
HOUSE ADYour CV gets thirty seconds.CV writing and honest review. English & Greek.kaeros.app →
AI Evaluation Engineer — Zafin · Job Opportunities API