Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

AI Evaluation Engineer

Judi Health
CompanyJudi Health
CategoryEngineering
LocationCharlotte
RemoteOn-site (inferred)
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted14 Jul 2026
Last verified30 Jul 2026
SourceEmployer career page (greenhouse)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
About Judi Health Judi Health is an enterprise health technology company providing a comprehensive suite of solutions for employers and health plans, including: Judi Rx , a public benefit corporation delivering full-service pharmacy benefit management (PBM) solutions to self-insured employers, Judi Health™ , which offers full-service health benefit management solutions to employers, TPAs, and health plans, and Judi® , the industry’s leading proprietary Enterprise Health Platform (EHP), which consolidates all claim administration-related workflows in one scalable, secure platform. Together with our clients, we’re rebuilding trust in healthcare in the U.S. and deploying the infrastructure we need for the care we deserve. To learn more, visit www.judi.health . Hybrid 3 days (offices in NYC, Denver, CO and Charlotte, NC area) Position Summary As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and autonomous agents in production. This role bridges the gap between model development and real‑world usage by translating ambiguous product goals into measurable quality targets. We’re looking for someone to lead evaluation end-to-end — from unit and integration testing to offline, online, and statistical evaluations of probabilistic systems. What we need is someone who can design and operate robust evaluation frameworks, partner with scientists and engineers, and ensure we can confidently answer questions like: “Did this change improve or degrade quality, safety, or user outcomes?” What You’ll Build Evaluation & Quality Pipelines Build data evaluation pipelines that collect production conversations and agent interactions Reconstruct full sessions from traces, logs, recordings, and transcripts Apply labeling and scoring using human feedback signals (surveys, sentiment, outcomes) and automated evaluators (e.g., LLM‑as‑judge) Continuous Quality & Safety Benchmarking Own weekly and on‑demand automated evaluation runs against staging and production Define benchmarks that track accuracy, reliability, and safety‑related signals Produce trend dashboards that clearly answer: “Did this deploy change quality or risk?”   Unified Evaluation Framework Design and extend a standardized evaluation framework that supports multiple agent types and workflows Translate high‑level product expectations into concrete success criteria and metrics Ensure new agents and features can be evaluated consistently with minimal friction Self Service Evaluation Tooling Build APIs and internal tools so data scientists and engineers can go from “interesting scenario” to “included in the eval suite” quickly Enable scenario curation, dataset management, and eval execution without deep infrastructure knowledge   Experiment Tracking & Visibility Provide shared visibility into prompt, model, and agent experiments Enable reproducibility and comparison across runs so teams can build on each other’s work instead of operating in silos Position Responsibilities: Data Engineering Build and maintain ETL pipelines for heterogeneous data sources (traces, logs, transcripts, user feedback) Implement complex data stitching and session reconstruction logic Manage dataset versioning, provenance, and lifecycle Platform & Observability Develop dashboards and monitoring tools for AI quality metrics Integrate evaluations into CI/CD pipelines for scheduled and gated runs Implement alerting on quality and safety signals, not just infrastructure health AI / ML Evaluation Tooling Apply and extend LLM‑as‑judge evaluation patterns Design metrics and scoring approaches suitable for stochastic, non‑deterministic systems Use tools like LangSmith to track runs, traces
HOUSE ADYou found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →
AI Evaluation Engineer — Judi Health · Job Opportunities API