Staff Data Engineer
Iterative Health
| Company | Iterative Health |
| Category | Engineering |
| Location | Cambridge |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 24 Apr 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (greenhouse) |
Description
Iterative Health is a healthcare technology and services company powering the acceleration of clinical research to transform patient outcomes.
We built a leading performance-driven network of 100+ sites across the US, Europe, India, and Australia, conducting research directly in the communities where care is delivered across gastrointestinal, hepatology, obesity, and cardiology. By combining deep clinical trial expertise with cutting-edge AI, we connect sponsors' scientific ambitions with high-performing research teams that expedite and expand access to novel therapeutics for patients in need. Today, Iterative Health is headquartered in Cambridge, Massachusetts, and New York City with 250+ employees world-wide.
About the Role
Accelerating clinical research is one of the defining challenges in healthcare. Promising therapies exist that patients can't access because the operational infrastructure to run clinical trials efficiently doesn't exist yet. We're building it. That means designing technology systems that bring order to a fragmented landscape of clinical data sources, automating the operational work that slows trials down, and turning real-world clinical data into a foundation for predictive intelligence.
We're building a uniquely valuable data asset: real-world patient and research data flowing across 80+ trial sites, spanning dozens of EHRs and clinical systems, focused on patient populations that are chronically underserved by existing clinical research infrastructure. Your job is to build the pipelines, data models, and AI infrastructure that make this asset real, from ingestion and normalization through to the systems that power predictions on top of it. You'll own data quality and observability as foundational engineering problems. You'll also have a direct hand in shaping how this data drives our AI strategy, what we model, what we predict, and what becomes possible.
This is an opportunity for someone who wants to be part of a small, fast-moving engineering team at a formative stage. You'll shape what gets built, how decisions get made, and what the team becomes.
Responsibilities
Own the data layer and architecture: the models, schemas, and infrastructure decisions that everything downstream depends on
Build and operate the pipelines and transformations that move data from ingestion through normalization, enrichment, and into the formats that support analytics, ML training, and production model serving
Own data quality and observability: build the systems that make data issues visible and correctable before they compound
Partner with ML and engineering teams to identify what's modelable, define training data requirements, and build the data foundations for new predictive capabilities
Define how clinical and operational data is governed across the system
Evaluate and select the tools and technologies that make up the data stack, with a clear point of view on build vs. buy
Help shape the engineering culture of a small, growing team: how technical decisions get made, how problems get debated, what rigor looks like in practice
What We’re Looking For
Required Qualifications
10+ years of experience in data engineering or related roles, with significant time spent building data systems
Experience with healthcare data strongly preferred (HL7, FHIR, claims, EHR extracts) or other complex, regulated data domains
Deep experience modeling and integrating data from multiple heterogeneous sources with inconsistent schemas and quality
Experience applying AI and LLMs to data engineering problems: extraction, normalization, classification, entity resolution
Strong understanding of how data infrastructure supports ML workflows from feature engineering to training data pipelines to model serving
Fluent in SQL and at least one modern programming language (Python, Java, Scala, Go), with experience across modern data infrastructure - distributed
You found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →