Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Senior Site Reliability Engineer

Nexla
CompanyNexla
CategoryEngineering
LocationBengaluru
RemoteOn-site (inferred)
EmploymentNot stated
LevelSenior
SalaryNot stated by the employer
Posted28 Apr 2026
Last verified3 Aug 2026
SourceEmployer career page (greenhouse)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
About Nexla Nexla is the leading Integration platform, built with AI, for AI. Nexla takes a metadata driven approach to converge diverse integrations across Data, Documents, Agents, Applications, and  APIs into a single design pattern. We accelerate the development of solutions for GenAI, Analytics, and Inter-company data. Nexla makes data users and developers up to 10x more productive by delivering a true blend of no-code, low-code, and pro-code interfaces. Leading companies including DoorDash, LinkedIn, Johnson & Johnson, and LiveRamp trust Nexla for mission-critical data. Named in the 2022, 2023, and 2024 Gartner Magic Quadrant™ for Data Integration Tools and top-rated by customers on Gartner Peer Insights, headquartered in San Mateo, California. At Nexla, our culture is built around our core values:  Have Empathy , Be Curious , Be Intellectually Honest , Achieve Excellence , and Remember to Relax . We put our customers at the heart of everything we do, foster a data-driven mindset, take ownership of our work, and believe in the power of teamwork to achieve ambitious goals. Role You will own the reliability of the distributed data systems at the heart of Nexla - the streaming runtime and processing engines that move hundreds of billions of rows per day for top-tier enterprises. This is an SRE role for our big data stack: Kafka, Spark, Flink, Ray, Redis, and data warehouses, all running on Kubernetes. This is not a cloud-provisioning role. We are looking for someone who has lived inside stateful, high-throughput systems in production who has chased down a broker outage, a checkpoint stall, a crashlooping cache, and a sink that silently stopped writing, and who fixes the architecture rather than the symptom. If keeping a large, busy data platform alive and fast is the kind of problem you find satisfying, you will have a lot of fun working with us. This is a unique opportunity to shape the foundation of a product that is defining the next wave of intelligent, context-aware data movement. Responsibilities Streaming & Data Plane Reliability: Own the health of our Kafka-based runtime (managed via Strimzi on Kubernetes) - broker health, topic lifecycle and count management, partition and throughput tuning, certificate/secret rotation, and version upgrades - at a scale of hundreds of thousands of topics and hundreds of billions of rows per day. Distributed Processing Engines: Operate and tune distributed system workloads in production in collaboration with backend teams, resource allocation, autoscaling, checkpointing, backpressure, and failure recovery for both batch and streaming jobs. Stateful Services: Run Redis clusters and other stateful systems reliably - failover, persistence, liveness/readiness tuning, and capacity planning under heavy and bursty load. Kubernetes & Operators: Take end-to-end ownership of Amazon EKS, Google GKE and the operators (Strimzi and others) running our stateful data workloads - cluster lifecycle, scaling, version upgrades, and resource governance. Observability: Build deep, data-aware monitoring - consumer lag, throughput, partition skew, job latency, error rates - not just host and CPU metrics. Make the data plane's behavior legible before it breaks. Incident Management: Lead root-cause analysis for distributed-systems failures (broker outages, crashloops, sink decommissions, control-plane race conditions) and drive durable fixes. Mitigate fast, but design out the recurrence. Infrastructure as Code & Automation: Provision and manage cloud infrastructure with Terraform; build operational runbooks and automation, including for air-gapped / private enterprise installs (pre-staged images, operator-facing procedures). Collaboration: Partner with platform, runtime, and connector engineering - and with SREs and support - to ship and scale new data-movement features reliably in a large-scale Linux environment.r with SREs,
HOUSE ADYour CV gets thirty seconds.CV writing and honest review. English & Greek.kaeros.app →