Senior Software Engineer, Agent Oversight
Scale AI
| Company | Scale AI |
| Category | Engineering |
| Location | San Francisco |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 14 Jul 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (greenhouse) |
Description
About Scale
Scale’s mission is to develop reliable AI systems for the world’s most important decisions. As the leading AI data foundry, we provide the high-quality data and full-stack technologies that power the world’s most advanced models — fueling breakthroughs in generative AI, defense, and autonomous vehicles. We partner with leading enterprises and governments to bring AI into production that performs when it matters most, combining rigorous evaluation with full-stack deployment so our customers can build AI they can trust.
About the Team
Applied Intelligence Systems team is part of the Scale Generative AI Platform (SGP), focused on pushing the frontier of what agentic applications can do across diverse enterprise and government use cases. We build the infrastructure and tooling that power Agentic AI in production, paired with applied ML research, design, and evaluation to ensure these systems perform reliably at the scale our customers demand. We’re growing fast, with increasing traction across both commercial and public sector customers, and we’re just getting started — this team will define what dependable, production-grade agentic AI looks like.
About the Role
As a Software Engineer on Agent Oversight, you will build the platform infrastructure that lets our production agents be observed, evaluated, and improved at scale. This includes building observability tooling, evaluation harnesses, and the pipelines that connect them to improvement loops. Whether building foundational infrastructure or partnering closely with ML engineers on production workflows, you will own your systems end-to-end while maintaining rigorous technical standards.
You will:
Design and build core platform capabilities for deploying, monitoring, and evaluating agentic applications in production
Build reliable APIs and data pipelines that capture agent telemetry, evaluation signals, and performance metrics at scale
Work alongside ML engineers where platform work intersects with evaluation or improvement systems — bringing enough ML fluency to reason about model behavior, evaluation quality, and improvement loops while owning the software systems that make those workflows reliable
Own the reliability, scalability, and observability of platform components serving multiple concurrent enterprise and government customers
Work cross-functionally with product, forward deployed engineering, and customers to translate real-world deployment requirements into platform features
Build features end-to-end: system design, implementation, debugging, and testing
Participate in high-velocity experimentation to validate platform capabilities against real customer usage
Requirements:
4+ years of professional software engineering experience, with strong fundamentals in backend/distributed systems, APIs, and data pipeline design
Hands-on experience building production software for ML/LLM-powered products or platforms, such as evaluation systems, observability/monitoring, experimentation infrastructure, agent runtimes, model-serving-adjacent services, or telemetry/data pipelines
Working knowledge of how LLM or ML systems behave in production: evaluation signals, failure modes, prompt/tool-calling workflows, experiment results, data quality issues, and the tradeoffs between offline evals and live customer behavior
Experience partnering closely with ML engineers or applied researchers to turn prototypes, eval loops, or model-improvement workflows into reliable platform capabilities, without needing to own model training, modeling strategy, or research direction
Experience building infrastructure or platforms that other engineering teams build on top of (internal platform, developer tools, or similar)
Track record of taking ownership of features or components end-to-end — from design through production — within a larger platform or system
Comfortable operating in an ambiguous, fast-changing domain where to
986,449 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →