AI Systems Architect | Remote
TubeScience-Labs
| Company | TubeScience-Labs |
| Category | Data & Analytics |
| Location | Remote |
| Remote | Remote |
| Employment | Not stated |
| Level | Lead |
| Salary | Not stated by the employer |
| Posted | 23 Apr 2026 |
| Last verified | 2 Aug 2026 |
| Source | Employer career page (greenhouse) |
Description
About TubeScience Labs
TubeScience Labs is the applied-AI team inside TubeScience — the largest performance-video company in paid social. TubeScience is Meta's largest creative partner and AppLovin's #1 creative partner, producing 8,000+ original ads every month from a 100,000 sq ft Los Angeles studio, backed by a library of 1.6 million+ performance ads and $2B in annual managed ad spend. That makes one of the richest first-party creative-performance datasets anywhere.
Labs turns that data — and the playbook behind billions in spend — into frontier AI tools that actually ship. Our products run in production against real creative, real deadlines, and real budgets every day. The tools that graduate internally become products we ship to external clients.
About the Role
TubeScience is hiring an AI Systems Architect to serve as the technical lead for the infrastructure at the core of our production AI platform — the media pipelines, LLM routing layers, distributed data systems, agentic AI systems, and observability tooling that everything else depends on. You'll own the technical direction for how these systems are designed, working closely with product and engineering across the org, and then you'll build them.
This is a deeply hands-on, greenfield role. You'll write code, set architectural direction, and make the hardest technical calls — but you won't be executing someone else's roadmap. You'll be one of a small number of engineers who decides not just how to build, but what to build, as the platform and the AI systems underneath it evolve in tandem. The decisions you make early will be load-bearing for years.
We care as much about craft as we do about capability. The right person obsesses over how systems behave under real load — reliability, observability, the failure modes that only show up in production — and knows how to put the foundations in place so a fast-moving team can keep that bar high.
This role is fully remote.
In this role, you will:
Own the infrastructure core : Set the technical strategy for the media pipelines, LLM routing, distributed data systems, and observability tooling that the entire platform runs on — and architect them from first principles.
Design AI agent systems : Build AI agents and multi-agent systems from the ground up — orchestration, tool-calling, MCP servers, A2A dispatch, and the permissions and observability infrastructure that makes them reliable in production.
Go deep on distributed systems : Push into the hard parts — data flow, fault boundaries, consistency trade-offs, performance under load — to build infrastructure that holds up when it matters.
Raise the bar on reliability : Design for observability from day one — tracing, metrics, structured logging — and build the quality systems that let a fast-moving team move without breaking things.
Shape the roadmap : Bring an engineering point of view to platform strategy — what we should build, in what order, and why.
You might be a good fit if you:
Have set technical direction across teams, and can point to specific systems you architected and operated in production, and the moments where you brought people along.
Are deeply expert in distributed systems including the layers below the framework — data flow, fault tolerance, consistency, and performance under real load.
Design AI agents and agentic systems: orchestration, tool-calling, multi-agent coordination, and the infrastructure that makes them observable and production-ready.
Have built large-scale platforms from scratch — making the foundational decisions and living with them — not just joined a mature codebase.
Design clean, durable APIs across REST, gRPC, event-driven, and schema-based systems, and are fluent with MCP, A2A, schema registries, and tool-calling normalization.
Know media pipelines deeply — transcoding, ffmpeg-class tooling, cloud storage, large-object proce