Senior Data Scientist - LLM Evaluation & Business Intelligence
paradigm-health
| Company | paradigm-health |
| Category | Data & Analytics |
| Location | US-Remote |
| Remote | Remote |
| Employment | Not stated |
| Level | Senior |
| Salary | USD 148k–184k |
| Posted | 23 Jul 2026 |
| Last verified | 11 Aug 2026 |
| Source | Employer ATS (ashby) |
Description
Paradigm Health is rebuilding the clinical research ecosystem by enabling equitable access to trials for all patients. Our platform enhances trial efficiency and reduces the barriers to participation for healthcare providers. Incubated by ARCH Venture Partners and backed by leading healthcare and life sciences investors, Paradigm’s seamless infrastructure implemented at healthcare provider organizations, will bring potentially life-saving therapies to patients faster.
Our team hails from a broad range of disciplines and is committed to the company’s mission to create equitable access to clinical trials for any patient, anywhere. Join us, and bring your expertise, passion, creativity, and drive as we work together to realize this mission.
ABOUT THE ROLE:
As a Data Scientist on our team, you will help define how we measure and trust the LLM systems that power Paradigm's clinical data products and build the business intelligence layer that turns that data into reliable reporting across the org. You'll work at the intersection of applied AI evaluation, analytics engineering, and clinical research, partnering with AI engineers, data engineers, clinicians, and non-technical stakeholders.
In this role you'll be responsible for delivering reliable data that helps Paradigm measure its business and drive outcomes. You'll partner with teams across the organization to define requirements, then research and build the models and reporting that meet them. This is a deeply collaborative, internally-facing role: your partners range from highly technical teammates to non-technical business, clinical, and product stakeholders, and much of your impact comes from translating between them — turning ambiguous questions into well-defined data products and communicating results, and their caveats, clearly to every audience.
Your primary focus will be evaluating production LLM pipelines: designing accuracy metrics, measuring retrieval effectiveness, and making principled cost/quality tradeoffs for the prompts and models that extract clinical meaning from patient data. Alongside this, you'll help establish a trustworthy business intelligence foundation — production data models, a shared semantic layer, and governed pipelines that give the whole organization consistent, defensible numbers. It's a role for someone who wants to bring scientific rigor to a fast-moving AI system and see their work drive measurable impact for patients and providers.
WHAT YOU'LL DO:
- Design and run evaluations of production LLM pipelines, developing accuracy and quality metrics that tell us how well our systems perform on real clinical tasks.
- Support evaluation of our Patient Trial Evaluation (PTE) pipeline — assessing prompts and trial-matching logic for accuracy in reporting, and surfacing where and why they fail.
- Measure and improve the effectiveness of our retrieval-augmented generation (RAG) systems, from retrieval quality through final output.
- Use accuracy metrics to inform cost/quality tradeoffs across trials, prompts, and models, giving the team a clear basis for which approaches to ship.
- Establish and maintain a library of "best prompts" for recurring clinical concepts, backed by evidence rather than intuition.
- Partner with AI Engineering to close the loop between evaluation findings and production improvements.
- Build and maintain production-grade data models using SQL and dbt on Databricks, ensuring analytics logic is reliable, maintainable, and production-ready.
- Help establish and grow a semantic layer that enables consistent, trusted reporting across the organization.
- Advance data governance practices across dbt, Databricks, and Hex — documentation, testing, versioning, and clear ownership of metrics.
- Collaborate with non-technical internal stakeholders to develop reporting logic, answer data questions, and translate business needs into well-defined data products.
- Support both internal and externa