Senior Software Engineer - AI Agent Platform
Navan
| Company | Navan |
| Category | Engineering |
| Location | Tel-Aviv |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 23 Jul 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (greenhouse) |
Description
Navan’s Cognition team builds and operates AI-powered travel and expense experiences used by customers around the world. We are developing intelligent agents that help users search, make decisions, and complete complex workflows through natural, personalized interactions.
Our team owns the AI platform end to end - from agent execution and LLM orchestration to state management, integrations, evaluation, observability, and production reliability.
We are looking for a Senior Software Engineer to own and evolve the runtime platform that powers Navan’s production AI agents.
This is a hands-on backend and platform engineering role for someone who combines deep TypeScript expertise, strong distributed-systems fundamentals, and exceptional production debugging skills with a practical understanding of LLMs and agentic systems.
You will work closely with engineers building AI agents. Your responsibility will be to provide the reliable runtime, infrastructure, abstractions, and observability they need to deliver new capabilities safely and quickly.
What you’ll do
Own and evolve the production runtime responsible for executing and orchestrating AI agents.
Design platform capabilities for agent execution, tool calling, streaming, state management, persistence, and long-running workflows.
Build resilient integrations with multiple LLM providers and model-serving platforms.
Design provider-routing and fallback strategies based on availability, latency, quality, and cost.
Implement retries, timeouts, circuit breakers, rate-limit handling, idempotency, and graceful degradation.
Ensure the platform remains available when external dependencies or infrastructure components experience outages.
Build reliable mechanisms for loading, caching, versioning, and recovering agent configurations and artifacts.
Create end-to-end observability for AI requests, including model, provider, agent, latency, token usage, cost, errors, retries, and fallback behavior.
Define dashboards, alerts, SLOs, and runbooks for production AI workloads.
Lead the investigation of complex production issues across application code, infrastructure, external providers, distributed state, and agent behavior.
Improve platform scalability, concurrency, latency, and resource efficiency.
Build reusable APIs and abstractions that allow agent developers to add capabilities without duplicating infrastructure logic.
Strengthen platform quality through integration testing, load testing, failure injection, and dependency-outage simulations.
Turn production incidents into architectural improvements, automated tests, monitoring, and operational safeguards.
Collaborate with product, infrastructure, and engineering teams to translate customer and business requirements into platform capabilities.
Mentor engineers and establish best practices for building and operating reliable production AI systems.
What we’re looking for
7+ years of professional software engineering experience, primarily in backend, platform, or distributed systems.
Expert-level TypeScript and Node.js skills.
Experience with NestJS or a comparable backend framework.
Proven experience designing, building, and operating large production services.
Strong understanding of distributed-systems patterns, including retries, backoff, idempotency, circuit breakers, caching, consistency, and failure recovery.
A systematic debugging mindset and the ability to trace failures across multiple services and dependencies.
Experience owning customer-facing systems where availability, latency, and correctness directly affect users.
Strong experience with cloud infrastructure and managed services, preferably AWS.
Experience with distributed caching and storage technologies such as Redis and S3.
Hands-on experience with production observability: structured logs, metrics, tracing, dashboards, alerts, and SLOs.
Experience participating in incident response and driving fo
982,094 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →