Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Senior Engineering Manager, Site Reliability

Horizon3ai
CompanyHorizon3ai
CategoryEngineering
LocationUS
RemoteRemote
EmploymentNot stated
LevelManager
SalaryNot stated by the employer
Posted16 Jul 2026
Last verified31 Jul 2026
SourceEmployer career page (ashby)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Get to Know Us Horizon3.ai is a fast-growing, remote cybersecurity company dedicated to the mission of enabling organizations to proactively find and fix and verify exploitable attack vectors before criminals exploit them. Our flagship product, the NodeZeroTM platform, delivers production-safe autonomous pentests and other key assessment operations that scale across the largest internal, external, cloud, and hybrid cloud environments. NodeZero has been adopted by organizations of all sizes, from small educational institutions to government agencies and Global 100 enterprises. It is used by ITOps/SecOps teams, consulting pentesters, and MSSPs and MSPs.  We are a fusion of former U.S. Special Operations cyber operators, startup engineers, and formerly frustrated cybersecurity practitioners. We're committed to helping solve our common security problems: ineffective security tools, false positives resulting in alert fatigue, blind spots, "checkbox” security culture, cybersecurity skills shortage, and the long lead time and expense of hiring outside consultants. Collectively, we are a team of learn it alls, committed to a culture of respect, collaboration, ownership, and results. We are looking for an engineering leader to build and lead the Site Reliability function at Horizon3. This leader is expected to hire a team and establish a Site Reliability function to define and drive investments in operational excellence to ensure Horizon3’s product and service offerings meet customer and business expectations. WHAT YOU’LL DO Establish an SRE function. - Build a SRE team, from scratch. Hire experienced site reliability staff and build a team of 4-6 in year one. - Professionalize incident management. Define and document incident processes and practices for your SRE team and for the application feature teams. Make tool and vendor decisions to support processes. - Drive incident professionalism and reliability culture across the engineering organization through training and process adoption. Drive Engineering Excellence - Use design reviews, code reviews, and blameless retrospectives to drive a culture of quality and excellence in engineering. People Leadership and Project Management - Balance incident response while also executing on a roadmap of observability and reliability engineering initiatives. - Hire and directly manage site reliability engineers. - As a Manager, you will be responsible for: - Recruiting and onboarding talented individuals to support our organizational goals - Mentoring, coaching, equipping, and developing your team - Recognizing and retaining high performers - Leading horizontally with peer Management & Senior Leaders WHAT YOU’LL BRING (QUALIFICATIONS) Building and Leading SRE Teams - Demonstrated experience leading hiring and growing SRE or Infrastructure teams. Experience leading or building SRE functions, including incident management processes, on-call programs, SLO/SLA definition, and operational runbooks. - Previous career experience as a Site Reliability Engineer. Comfortable in being hands-on while you grow and hire your team. - Deep hands on experience with observability: application performance management, logs and traces, and golden signals and service-specific metrics. - Experience in selecting and deploying incident management tooling (e.g., PagerDuty, FireHydrant,etc.) and creating decision making frameworks for vendor selection. - Strong working knowledge of at least one major cloud provider (AWS, GCP, or Azure) — infrastructure, networking, managed services, IAM, cost management. AWS preferred. Engineering Culture and Process - Experience in design practices like architecture decision records or RFCs in the infrastructure or platform engineering context. - Able to write clear, durable process documentation that engineering teams can adopt. - Experience in providing education and training on runbook development, o
HOUSE AD1,193,877 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →
Senior Engineering Manager, Site Reliability — Horizon3ai · Job Opportunities API