Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Principal Engineer, Cloud Infrastructure

Clover Health
CompanyClover Health
CategoryEngineering
LocationRemote - USA
RemoteRemote
EmploymentNot stated
LevelLead
SalaryNot stated by the employer
Posted22 May 2026
Last verified30 Jul 2026
SourceEmployer career page (greenhouse)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
We're building the cloud infrastructure and delivery platform that one of healthcare's most ambitious technology transformations runs on, and we need a Principal Engineer, Cloud Infrastructure to own the infrastructure behind it all. You'll both set the technical direction and do the work hands-on, while leading a small team of infrastructure engineers. This role is about infrastructure and developer enablement. As our AI strategy evolves, we see a world where many outside of traditional engineering and data roles will contribute to our technology ecosystem. We believe reliability, security, and cost discipline come from good platform design and automation, and that ticket queues and runbooks are signs we haven't solved the underlying problem. You'll bring a clear point of view on what a modern, AI-native infrastructure org should look like, the judgment and technical skill to make it real, and the leadership instincts to build and develop the contributors around you. In this role you will: Own cloud infrastructure end-to-end across our cloud environments, staying hands-on with IaC, CI/CD, and observability alongside the team Own the SDLC surface area (dev environments, lower envs, release mechanics, monitoring and response) so engineers can ship in minutes rather than days Build the substrate that lets AI agents observe, reason, and act on business systems with safe defaults, and the rails that let AI-assisted apps from people outside of tech be deployed and supported without becoming operational liabilities Participate in security operations alongside our SecOps team on CSPM, threat management, and patch management as the infrastructure-side owner of remediations Frame infrastructure decisions in terms of business outcomes (operational efficiency, financial impact, clinical results), and make progress and tradeoffs visible to leadership as a natural byproduct of how you work Lead and develop a small, high-impact team, and operate as a peer to our SecOps, Data, IT Systems, and App Engineering leaders as a unifying technical force across the org You should get in touch if: You have an engineering background with deep, hands-on experience operating production infrastructure across at least two major clouds (GCP, AWS, Azure) and with modern infrastructure-as-code (Terraform or similar) You have built and operated modern CI/CD systems, dev environments, and observability stacks. What you've shipped matters more than the size of the fleet you've managed. You've built and operated infrastructure across different company sizes and business contexts. Cross-domain breadth is a real asset; experience in regulated domains (healthcare, financial services) is helpful but not required You hold a strong, well-reasoned point of view on platform philosophy, and can defend when to standardize, when to give teams room, and how to make the safe path the easy path You have shipped AI-native operational systems. That might be AI agents that triage alerts, draft RCAs, manage cost, or execute runbooks, or frameworks that let non-traditional contributors (junior engineers, analysts, AI agents, vibe-coders) ship safely. You're action-oriented on governance and security participation. You document what matters, automate what you can, and don't let process become a bottleneck You have built and developed teams. You think about composition intentionally, hire for gaps, and invest in people's growth You're a strong communicator: clear, direct, and concise with both technical and non-technical audiences Success in this role looks like First 90 Days: Develop a clear point of view on the current state of the infrastructure, CI/CD, and operational tooling. Identify the highest-leverage opportunities. Ship a first step-function improvement or eliminate a fundamental operational limitation. 6 Months: AI-assisted operations are in place for alert triage and incident respon
HOUSE AD995,367 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →