Director, Site Reliability Engineering
Stellar
| Company | Stellar |
| Category | Engineering |
| Location | New York |
| Remote | Hybrid |
| Employment | Not stated |
| Level | Director |
| Salary | Not stated by the employer |
| Posted | 4 Jun 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (ashby) |
Description
Interested in working on cutting-edge blockchain technology and creating equitable access to the global financial system? Since 2014, the mission-driven team at the Stellar Development Foundation (SDF) has helped fuel the tremendous growth of the Stellar blockchain network, an open-source platform that operates at high-scale today. Developers and companies around the world build on it, and the SDF team is expanding to support the rapidly growing and changing Stellar ecosystem.
SDF is looking for a Director of Site Reliability Engineering to lead a small, high-leverage SRE team and help shape how engineering teams own, operate, and improve production services.
This is a senior engineering leadership role reporting to the CTO. You will set the vision, operating model, and culture for SRE while owning the core infrastructure services that help SDF engineering teams build, deploy, observe, and operate software with confidence.
Engineering teams at SDF own the services they build. SRE provides the frameworks, standards, shared infrastructure, tooling, observability practices, and enablement model that make strong service ownership possible across engineering.
You will be successful here if you bring strong technical judgment, pragmatic leadership, and the ability to influence through trust, clarity, and execution. SDF is a small, mission-driven foundation with a broad technical surface area, so this role requires leverage, ownership, and a bias toward solving the right problems over creating processes for its own sake.
In this role, you will:
- Lead, coach, and develop a distributed SRE team, setting a clear vision, charter, operating model, priorities, and success measures.
- Define and roll out a Service Ownership & Maturity Framework across engineering, with expectations that vary appropriately by service criticality.
- Own and improve core engineering infrastructure services, including cloud foundations, Kubernetes and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation.
- Help engineering teams become stronger owners and operators of their services through better standards, dashboards, runbooks, alerting, escalation paths, operational readiness, and deployment practices.
- Make reliability, operational maturity, infrastructure health, and developer productivity more measurable through trusted metrics and practical operational intelligence.
- Improve deployment automation, resilience, self-healing patterns, disaster recovery readiness, and service reliability based on actual impact and risk.
- Mature incident response, escalation, postmortems, and on-call health across a geographically distributed team.
- Build paved paths and self-service infrastructure that reduce toil, lower cognitive load, and help engineering teams move faster while strengthening ownership and reliability.
- Partner closely with Security, Compliance, Legal, Finance, Procurement, and Corporate IT where infrastructure, access management, cloud operations, vendor review, or controls intersect with engineering.
- Pragmatically evaluate AI-assisted and agentic workflows where they can improve infrastructure operations, service ownership, developer workflows, or toil reduction.
You have:
- 10+ years of experience in SRE, infrastructure engineering, platform engineering, cloud infrastructure, production operations, or closely related engineering roles.
- 5+ years of experience leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers.
- Strong experience defining team charters, operating models, roadmaps, success measures, and engineering practices for infrastructure or reliability teams.
- Deep technical judgment across cloud infrastructure, production operations, distributed systems, reliability tradeoffs, automation, and operational risk.
- 3+ years of experience with modern cloud infrastructure in AWS, GCP
You found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →