Senior ML Research Scientist, Perception Models
Twelve Labs
| Company | Twelve Labs |
| Category | Data & Analytics |
| Location | Seoul |
| Remote | Hybrid |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 20 Jul 2026 |
| Last verified | 8 Aug 2026 |
| Source | Employer ATS (ashby) |
Description
WHO WE ARE
Video is 90% of the world's data. Most of it is invisible to machines.
TwelveLabs builds the intelligence layer to change that. Our multimodal AI models understand video the way humans do — across sight, sound, and motion — and power production-scale AI workloads across media, entertainment, sports, security, and government.
We have raised more than $210 million from NEA, Radical Ventures, Amazon, NVIDIA, Snowflake, Databricks, Index Ventures, NAVER Ventures, Korea Investment Partners, Quadrille Capital, Red Bull Ventures, and AI pioneers including Fei-Fei Li, Silvio Savarese, and Alexandr Wang.
We are a global company, headquartered in San Francisco with offices in Seoul, New York, and London, and employees around the world. We believe the differences in our cultural, educational, and life experiences make our products stronger. Building technology that understands the world in all its complexity requires people who see it from every angle. We are looking for individuals who are driven by hard problems and want their work to matter. Come build it with us!
ABOUT THE TEAM
TwelveLabs builds multimodal foundation models that understand what happens in video and how visual information is organized across time and space. The Perception Models team develops general-purpose representations that work across full videos, clips, regions, and individual entities.
Rather than building a collection of narrow, task-specific models, we use signals from capabilities such as detection, segmentation, and tracking to strengthen shared video representations.
ABOUT THE ROLE
As a Senior ML Research Scientist, you will lead research at the intersection of spatiotemporal modeling and multimodal representation learning. You will define open-ended research problems, validate ideas through rigorous large-scale experiments, and turn promising results into production capabilities.
This role is ideal for a hands-on researcher with deep expertise in at least one of these areas and a strong interest in connecting them.
IN THIS ROLE, YOU WILL
- Research multimodal representations that connect temporal and spatial structure with semantic understanding.
- Develop models that operate across multiple levels of granularity, from full videos and clips to regions and entities.
- Design training objectives, datasets, evaluation methods, and large-scale experiments.
- Use task-specific vision models as supervision or components, and integrate useful signals into general-purpose representations.
- Partner with research and engineering teams to scale and ship new capabilities.
Even if you don't check every box, we encourage you to apply.
If you're a zero-to-one achiever, a ferocious learner, and a kind team player who motivates others, you'll find a home at TwelveLabs.
YOU MAY BE A GOOD FIT IF YOU HAVE
- A strong research track record in computer vision, video understanding, multimodal learning, or a related field.
- Deep expertise in at least one of the following: video foundation models, self-supervised or contrastive learning, embeddings, and retrieval; or temporal modeling, object-centric learning, detection, segmentation, and tracking.
- Strong hands-on skills in Python and PyTorch or an equivalent deep learning framework.
- Proven ability to design and run rigorous experiments at scale.
- Research impact through publications, production systems, or both.
- The ability to own ambiguous problems and collaborate across research and engineering.
We evaluate based on relevant technical skills and sustained industry impact. This role is typically a strong fit for engineers with an MS and deep industry experience who have evolved from individual contributor to technical leader in production ML environments.
PREFERRED QUALIFICATIONS
- Experience with open-vocabulary vision or dense and region-level representations.
- Experience with large-scale video training, data curation, or evaluatio