Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Software Engineer II - Platform & Infrastructure

Abnormal
CompanyAbnormal
CategoryEngineering
LocationHybrid - Bangalore
RemoteHybrid
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted26 May 2026
Last verified30 Jul 2026
SourceEmployer career page (greenhouse)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
About the Role Abnormal AI is looking for an experienced and driven Platform & Infra software engineer to join the PI team. Join us and help build the platforms that power Abnormal's growth Observability Platform - Own and evolve the monitoring, metrics, and alerting infrastructure that every engineering team at Abnormal depends on. You'll work across the Prometheus, Chronosphere, and Grafana stack to ensure engineers can see what their systems are doing in real time — building dashboards, managing metric pipelines at scale, operating the PagerDuty alerting pipeline, and driving cost-efficient observability across all production environments (US, EU, and GovCloud).   Your Impact Own the observability stack (Prometheus, Chronosphere, Grafana, PagerDuty) that every team relies on to detect, diagnose, and resolve production issues — when you make it better, every engineer at Abnormal gets faster. Design platforms and developer tooling that remove friction — reducing deployment times, simplifying pipeline authoring, and letting product teams focus on building rather than firefighting. Drive SLAs and SLOs for critical shared infrastructure ensuring the systems behind our products are resilient and cost-efficient.  Your architectural decisions on alerting pipelines and cross-environment deployments will define what products we can build and how quickly we deliver them to customers. What you will do  Work with the Tech Lead, Engineering Manager, and Product Manager to design, develop, and deliver key platform features — from technical design docs through production rollout Own features end-to-end: scoping, implementation, testing, deployment, and post-launch monitoring across multiple environments (US, EU, GovCloud) Take ownership of 1-3 key services within Observability (Prometheus, Chronosphere, Grafana, PagerDuty pipeline) or Data Infra (Airflow, Spark) and be accountable for their reliability, performance, and evolution Participate in on-call rotations — triage, diagnose, and resolve production issues independently, building deep operational knowledge of the systems you own Improve system resilience by converting runbooks into automated solutions, refining SLAs/SLOs, and proactively identifying performance bottlenecks and failure modes Assume ownership of the reliability of everything you build, including comprehensive unit tests, integration testing, and observability instrumentation Build platforms, tooling, and APIs that make it easier for other engineering teams to ship — whether that's faster pipeline deployments, better dashboards, or simpler alerting configuration Partner with internal customers (product and engineering teams) to understand their needs and translate them into scalable platform capabilities Communicate effectively in an async-first, distributed environment — proactively providing updates, discussing challenges, and proposing solutions without prompting Mentor junior engineers on the team, helping them ramp up on service operations and development practices Raise the bar of engineering excellence through code reviews, knowledge sharing, design discussions, and contributing to team best practices Must Haves  Backend Engineering & Distributed Systems (4+ years) 4+ years of hands-on backend engineering experience designing, building, and operating production-grade distributed systems Strong proficiency in Python — the primary language for Airflow DAGs, platform services, and automation tooling Working proficiency in Golang — used for high-performance infrastructure components, metric pipelines, and platform services Experience building systems that process data at scale — whether metric ingestion pipelines, stream/batch processing, or high-throughput API services Demonstrated experience owning a service or platform end-to-end — from technical design through production deployment, mo
HOUSE AD995,367 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →