Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Machine Learning Infrastructure Engineer

Universalagi
CompanyUniversalagi
CategoryEngineering
LocationSan Francisco
RemoteOn-site (inferred)
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted26 Feb 2026
Last verified30 Jul 2026
SourceEmployer career page (ashby)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
📍 San Francisco | Work Directly with CEO & founding team | Report to CEO | OpenAI for Physics | 🏢 5 Days Onsite MACHINE LEARNING INFRASTRUCTURE ENGINEER Location: Onsite in San Francisco Compensation: Competitive Salary + Equity Who We Are UniversalAGI is building OpenAI for Physics. AI startup based in San Francisco and backed by Elad Gil (#1 Solo VC), Eric Schmidt (former Google CEO), Prith Banerjee (ANSYS CTO), Ion Stoica (Databricks Founder), Jared Kushner (former Senior Advisor to the President), David Patterson (Turing Award Winner), and Luis Videgaray (former Foreign and Finance Minister of Mexico). We're building foundation AI models for physics that enable end-to-end industrial automation from initial design through optimization, validation, and production. We're building a high-velocity team of relentless researchers and engineers that will define the next generation of AI for industrial engineering. If you're passionate about AI, physics, or the future of industrial innovation, we want to hear from you. About the Role UniversalAGI is hiring an Infrastructure Engineer to build and own the execution platform powering our research and customer deployments: data generation + simulation orchestration + training/fine-tuning infrastructure + benchmarking pipelines + production deployments in customer environments. You’ll work closely with the CEO and founding team to turn research into repeatable, scalable, reliable systems - internally and in customer infrastructure. This is a “ship outcomes” role: your work directly determines how fast we can iterate, how reproducible our results are, and how reliably we deliver in production. What You’ll Do BUILD THE FOUNDATION PLATFORM (INTERNAL) - Build and operate scalable infrastructure for data generation and simulation workflows (job orchestration, scheduling, queues, retries, observability). - Build reproducible pipelines for training/fine-tuning and benchmarking (artifact/version management, experiment tracking, dataset lineage). - Own cost/performance tradeoffs across compute, storage, networking, and runtime efficiency. DEPLOY TO CUSTOMERS (EXTERNAL) - Lead deployments of our stack into customer cloud/on-prem environments, including secure networking, permissions, and data movement. - Build robust deployment patterns: environment provisioning, CI/CD, rollbacks, monitoring, and incident response. - Partner with customers to ensure reliability and repeatability under real-world constraints (security, compliance, infra limits, data governance). Qualifications - Strong software engineering skills (clean code, debugging, reliability, reproducibility). - Hands-on experience building/operating infrastructure for ML/compute-heavy workflows: pipelines, job orchestration, GPU compute, storage, CI/CD, monitoring. - Olympic athlete mindset: You have high standards for yourself and are obsessed with measurable improvement on the metrics you are delivering to customers. - Resourcefulness: you know when to do the “quick & correct” fix vs. when to invest in a robust solution, and you can justify the tradeoff with impact/ - Ownership: Comfortable owning work end-to-end and being accountable for measurable outcomes. Bonus Qualifications - Experience with workflow orchestration (e.g., Ray, Kubernetes, Slurm). - Experience with GPU infrastructure and distributed training systems. - Experience building evaluation/benchmarking frameworks with strong reproducibility guarantees. - Experience deploying into regulated / security-sensitive environments (gov/defense/enterprise). - Experience with simulation/HPC pipelines (CFD, meshing, batch workloads) is a plus but not required. - Experience in an FDE-style / delivery execution role (or similar “ship results fast” environments). Cultural Fit - Technical Respect: Ability to earn respect through hands-on technical contribution - Intensity: Thrives in our unusually
HOUSE ADYou found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →