Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Machine Learning Engineer, Speech - Joint Audio-Video Modeling

Cantina
CompanyCantina
CategoryUncategorised
LocationRemote (U.S. or Europe)
RemoteRemote
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted30 Jul 2026
Last verified2 Aug 2026
SourceEmployer ATS (ashby)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
About Cantina: Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning. If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!   About the Role: We're looking for a Research / ML Engineer to join our Speech Team to build state-of-the-art speech and audio generation systems end-to-end from data specs through production inference with a focus on joint audio-video modeling. You'll own the audio side of multimodal generation: the representations (audio VAEs, neural codecs), the generative backbone (diffusion / flow-matching transformers), and the conditioning and alignment machinery that makes characters speak, sing, and emote in sync with what's on screen. That includes voice cloning and multi-speaker conditioning inside joint AV models, cinematic dialogue with music and sound design, and adjacent speech tasks (controllable TTS, voice conversion) that feed the same stack. You'll drive the model ↔ data ↔ eval flywheel, partnering closely with research, video, data, and infra to ship fast, reliable, and cost-aware models. In this role you'll work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems. You will thrive in this role if you: - See research and engineering as two sides of the same coin and enjoy owning work end-to-end. - Are excited to work across modalities and collaborate closely with a video generation team rather than staying inside audio. - Are results-oriented, flexible, and willing to pick up whatever moves the needle. - Like collaborating closely with infra, data, and product to ship measurable improvements. - Enjoy designing experiments, listening tests, and metrics that correlate with user-perceived quality. - Are eager to learn every day, and to find and solve unique large-scale problems.   What You’ll Do: - Audio Representations: Design, train, and improve the audio VAEs, neural codecs, and vocoders our generative models sit on top of latent design, reconstruction and perceptual objectives, compression-vs-fidelity tradeoffs. - Model Building: Architect, implement, pre-train, fine-tune, and post-train/alignment (e.g., GRPO/DPO) diffusion and flow-matching transformers for large-scale audio and video generation. - Joint Audio-Video Modeling: Design the audio conditioning and cross-modal alignment inside joint AV models, audio latents alongside video latents, reference-audio and multi-speaker conditioning, multi shot generation audio/video modeling. - Experimental Design: Design, run, and analyze scientific experiments to advance our understanding of the models. - Data Ownership: Define data requirements and collaborate on acquisition, curation, AV-sync and quality filtering, annotation quality, and synthetic data strategies for paired audio-video and speech corpora. - Rigorous Evaluation: Design automated objective/subjective evaluations audio fidelity and intelligibility metrics, AV-sync, listening and viewing tests, robustness & bias checks, and red-team studies. - Inference Efficiency: Drive distillation, step-count reduction, quantization, and kernel/memory optimization to meet interactive latency and cost targets. - Pipeline Delivery: Harden the training → evaluation → inference pipeline; profile latency, memory, and cost; and meet production SLAs with robust monitoring and rollback. - GPU Scaling: Partner with infrastructure to run distributed training/inference on cloud fleets and productionize models with reliability and observability. - Project Leadership: Independently lead small research projects while colla
HOUSE ADYou have the idea. We build it.