Technical Product Manager, Data Ingestion & Quality
Protege
| Company | Protege |
| Category | Product |
| Location | Remote |
| Remote | Remote |
| Employment | Not stated |
| Level | Manager |
| Salary | Not stated by the employer |
| Posted | 7 Jul 2026 |
| Last verified | 6 Aug 2026 |
| Source | Employer ATS (ashby) |
Description
Company Overview:
We are building Protege to solve the biggest unmet need in AI — getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.
Solving AI’s data problem is a generational opportunity. We’re backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI. The company that succeeds will be one of the largest in AI — and in tech.
We’re a lean, fast-moving, high-trust team of builders who are obsessed with velocity and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.
About Protege
We are building Protege to solve the biggest unmet need in AI — getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.
Solving AI’s data problem is a generational opportunity. We’re backed by world-class investors and already powering partnerships with some of the most ambitious teams in AI. The company that succeeds will be one of the largest in AI — and in tech.
We’re a lean, fast-moving, high-trust team of builders who are obsessed with velocity and impact. Our culture is built for people who thrive on ambiguity, own outcomes, and want to shape the future of data and AI.
Role Overview
We're hiring a Product Manager to own the supply side of Protege's data platform — the pipeline that takes raw data from a partner and turns it into something catalog-ready, trustworthy, and usable. Right now, that process is manual, inconsistently applied, and a source of delivery risk. Your job is to change that.
This is a horizontal platform role, not a vertical one. You own the infrastructure that makes data trustworthy enough to build products from in the first place: the validation gates, the metadata generation pipelines, the QA standards, the de-identification transformations, and the catalog-readiness criteria that let the rest of the organization actually trust what's in our catalog.
You'll work across healthcare, media, and any other vertical we enter. You'll write SQL, review pipeline outputs, define what "good" looks like at each stage of ingestion, and translate those standards into platform requirements that engineering can build against.
The supply side is where data quality is won or lost. If this layer isn't working, nothing downstream works. It's foundational, largely invisible to customers, and one of the most important things we can build.
What you'll work on
Ingestion pipeline product: define the stages, validation gates, and quality checks that data passes through from partner arrival to catalog-ready; own the platform requirements that make this repeatable across modalities and verticals
Metadata generation: own the product decisions around what metadata gets extracted or generated at ingestion, including transcripts, tags, confidence scores, schema inference, at what threshold, and how it gets stored and surfaced
QA standards and tooling: define what "catalog-ready" means, build the tooling that enforces it, and get into the data directly to validate that standards are being met; you’ll run queries and review pipeline outputs, not just read dashboards
Cross-vertical consistency: work with vertical stakeholders to translate their "what does ready mean for our vertical" requirements into consistent platform-level standards that don’t require custom engineering per deal
What Success Looks Like
30 days: Ramp: Build a clear understanding of Protege’s current data ingestion workflow, including how raw partner data moves from arrival to catalog-ready. Get hands-on with pipeline outputs, schemas, metadata, validation check