Senior Backend Engineer
Judgmentlabs
| Company | Judgmentlabs |
| Category | Engineering |
| Location | San Francisco |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 15 Jun 2026 |
| Last verified | 31 Jul 2026 |
| Source | Employer career page (ashby) |
Description
SENIOR BACKEND ENGINEER
San Francisco · On Site · Full Time
Judgment Labs is building the infrastructure for continual learning in long-horizon AI agents.
The next generation of agents will not improve from prompts alone. They will improve from experience: the tasks they attempt, the tools they use, the mistakes they make, the edge cases they encounter, and the outcomes they produce in production. The hard part is turning that raw experience into high-quality data that can actually improve the system.
Judgment builds the infrastructure to do that. We turn long agent trajectories into clean, structured data for evals, labeling, rubric generation, context engineering, and RL workflows. Instead of only showing teams what happened, Judgment helps decide what matters, what should be learned from, and how that learning should flow back into the agent.
Databricks built the data infrastructure for analytics. Judgment is building the learning infrastructure for agents.
We’ve raised $30M+ from Lightspeed, SV Angel, Valor Equity Partners, and others.
THE ROLE
We’re looking for a Senior Backend Engineer to own the systems that ingest, structure, evaluate, and serve agent experience data at production scale.
This role includes the backend and data infrastructure surface area: high-throughput telemetry ingestion, ClickHouse-backed OLAP performance, evaluation pipelines, RabbitMQ/Temporal workflows, multi-tenant scheduling, and product-facing APIs. Some weeks you’ll be deep in distributed systems and query performance. Other weeks you’ll ship a customer-facing feature end to end across backend, frontend, and the data layer.
This is not a narrow API role. The backend is where raw agent trajectories become structured learning data.
INTERESTING TECHNICAL CHALLENGES
- High-throughput telemetry ingestion. Parse and persist OTEL traces across protobuf and JSON formats at hundreds of thousands of spans per second, writing to ClickHouse while keeping ingest latency low and backpressure graceful as customer traffic spikes.
- Petabyte-scale OLAP performance. Design schemas, partitioning, indexes, storage layouts, and query paths so behavioral queries over billions of spans stay fast. Turn real access-pattern analysis into concrete data modeling decisions.
- Long-horizon trajectory modeling. Agent workflows are messy: multi-step tasks, tool calls, retries, partial failures, context changes, and unclear outcomes. Build the abstractions that turn those trajectories into structured data for evals, labeling, rubric generation, context engineering, and RL workflows.
- Queue- and workflow-driven evaluation. Evaluations fan out across RabbitMQ and Temporal workflows. Getting this right means reasoning about retries, timeouts, idempotent state transitions, exactly-once-ish semantics, and reconciling runs that fail partway so nothing is silently orphaned.
- Multi-tenant fairness at scale. A single large customer should not be able to starve everyone else. Build scheduling and execution systems so latency stays predictable across hundreds of teams sharing the same evaluation pipeline.
- Near-real-time scoring. Behavioral scorers and agent judges call LLM APIs at scale, so batching, rate-limit management, retry/backoff, failure handling, and cost control are core backend systems problems.
- Learning loops for agents. Build the product and systems layer that helps teams decide what matters, what should be learned from, and how that learning flows back into the agent.
WHAT YOU’LL DO
- Design and build backend systems for trace ingestion, trajectory processing, evaluation orchestration, scoring, labeling, rubric generation, and customer-facing analytics.
- Own the API surface used by the Judgment platform UI, SDKs, JudgmentHub libraries, MCP server, Slack agent, and customer integrations.
- Build and operate the RabbitMQ / Temporal evaluation pipeline, including retry semantics, failure recovery, state reconciliation, a
You found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →