Senior ML Infrastructure Engineer
Prior Labs
| Company | Prior Labs |
| Category | Engineering |
| Location | Berlin |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 25 Jun 2026 |
| Last verified | 31 Jul 2026 |
| Source | Employer career page (ashby) |
Description
WHO WE ARE
Foundation models transformed text and images. Structured data - the largest and most consequential data format in the world - stayed untouched. Tables run every clinical trial, every financial model, every scientific experiment, every business decision, and no one had built a foundation model that truly understood them.
Until now. What LLMs did for language, we're doing for tables. The next modality shift in AI is happening, and we're hiring the team that makes it.
Momentum. We pioneered tabular foundation models and are now the world-leading organization in structured-data ML. Our TabPFN v2 model was published as a Nature https://www.nature.com/articles/s41586-024-08328-6 cover story and set a new state of the art for tabular machine learning. Since release we've scaled model capabilities 20x+, passed 3.5M+ downloads and 7,500+ GitHub stars, and are seeing accelerating adoption across research and industry - from detecting lung disease with Oxford Cancer Analytics https://www.oxcan.org/news/prior-labs-and-oxford-cancer-analytics-partner-to-advance-liquid-biopsy-and-clinical-decision-making-in-lung-disease to preventing train failures with Hitachi https://siliconangle.com/2025/12/01/prior-labs-debuts-tabular-ai-foundation-model-scales-10-million-rows/ to improving clinical-trial decisions with BostonGene https://priorlabs.ai/case-studies/boston-gene.
The hardest work is ahead. We're scaling tabular foundation models to millions of rows, thousands of features, real-time inference, and entirely new data modalities, while building the infrastructure to run them in production across some of the most demanding industries on earth. These are open problems no one else is working on at this level.
Our team. We're a small, highly selective team https://priorlabs.ai/about of 30+ engineers, researchers, and GTM specialists, with backgrounds spanning Google, Apple, Amazon, DeepMind, Meta, Microsoft Research, G-Research, Jane Street, Goldman Sachs, and CERN. We're led by Frank Hutter https://www.linkedin.com/in/frank-hutter-9190b24b/, Noah Hollmann https://www.linkedin.com/in/noah-hollmann-668b9010b/, and Sauraj Gambhir https://www.linkedin.com/in/sauraj-g/, and advised by world-leading AI researchers including Bernhard Schölkopf and Turing Award winner Yann LeCun. We ship fast, do top-tier research, and hold each other to an extremely high bar.
What's next. In 2025 we raised €9m pre-seed led by Balderton Capital, backed by leaders from Hugging Face, DeepMind, and Black Forest Labs. The next phase of growth is here, which makes this an ideal time to join.
ABOUT THE ROLE
We spend tens of millions per year on GPU compute to train tabular foundation models. That's not a target, it's what we're running today, and it's growing. The person who owns this infrastructure makes decisions worth millions of dollars: cluster architecture, scheduling efficiency, provider strategy, hardware selection. A wrong call costs six figures.
Today we run Slurm on GCP across multiple clusters. We're scaling to multi-cluster, multi-provider infrastructure and evaluating new hardware generations as they come online. You own the full stack, from cluster operations and cost optimization to distributed training performance and the tooling layer that keeps researchers moving fast. You work directly with the research team and understand what they're doing well enough to make infrastructure decisions that actually help them. And this isn't a pure support role. We operate an open environment. If you've got the next SOTA tabular architecture up your sleeve, go ahead and train it.
What you'll work on:
- Own and evolve multi-cluster GPU infrastructure. Slurm on GCP today, multi-provider and new hardware tomorrow. Architecture, scheduling, reliability, cost optimization
- Drive GPU utilization and training throughput: profiling, memory optimization, communication bottlenecks, systems-level debugging of distributed training across large runs
1,014,484 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →