Senior Platform Engineer
Element451
| Company | Element451 |
| Category | Engineering |
| Location | Remote |
| Remote | Remote |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 2 Jul 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (ashby) |
Description
Element451 is building the AI-powered platform reshaping how colleges and universities recruit, enroll, and support their students - and the reliability of that platform is what earns the trust of the institutions that run on it. This is a rare chance to own reliability, operations, and security at the level where it actually happens: keeping a real production system healthy, fast, and safe, and building the automation that keeps it that way. If you're a broad operator who's happiest when a single week spans operations, delivery, security, infrastructure, and data - and you've done that work where things move fast and the edges go unowned - you'll feel at home here.
THE ROLE
You'll keep Element451's platform reliable, secure, and operable - and build the delivery systems that let it scale without scaling the firefighting. This is a hands-on senior IC role, and a deliberately broad one: reliability and operations are the core, with CI/CD and delivery, security, infrastructure, and data reliability all real, recurring parts of the work. You'll partner closely with our Director of Platform Engineering, who owns the platform strategy; you own a large share of the operational and delivery execution that brings it to life.
We hold a high bar - and we give you the ownership, context, and support to meet it. The work has range and a fair amount of unpredictability; the people who thrive in it tend to want exactly that.
WHAT YOU'LL OWN
You own the operational health of the platform - that it's available, fast, observable, and safe in production - and you build the automation that makes those properties durable rather than heroic. In practice, you:
- Own the reliability discipline in practice - define and track SLIs and SLOs, keep the observability stack sharp (we use CloudWatch, Sentry, Papertrail, and Langsmith), and make system health and customer impact legible in real time.
- Carry production operations day to day - participate in on-call, lead incident response, run blameless post-incident reviews, and drive issues to root cause rather than patching symptoms.
- Treat operational toil as engineering work to eliminate - relentlessly automate remediation, sharpen alert quality, and drive down MTTD and MTTR rather than absorbing manual load.
- Own and evolve the CI/CD and delivery platform alongside the Director of Platform Engineering - build pipelines, deployment automation, environment management, and release tooling - so shipping is routine, safe, and low-drama, with progressive delivery, automated rollback, and production validation gates as standard.
- Build the developer-facing automation and paved roads that cut friction for the product engineering team - treating them as the platform's customer and making the reliable, secure path the easy path.
- Be the platform's hands-on security operator - IAM and least-privilege hygiene, secrets management, threat detection and response (WAF, GuardDuty), and vulnerability triage and remediation against SLA. This is the security function in practice today; you partner with the Director of Platform Engineering on strategy and standards, and you're trusted to set the operational bar where none exists yet.
- Keep the platform audit-ready by default - produce the infrastructure and reliability evidence that SOC 2 Type II and FERPA obligations depend on (we use Vanta), so audits are a byproduct of good operations, not a scramble.
- Build and operate cloud infrastructure as code - AWS (ECS/Fargate, Lambda, SQS/SNS, EventBridge, S3/CloudFront, VPC) managed with Terraform, with no manual or snowflake infrastructure in any environment.
- Plan and execute scaling ahead of product growth, and keep operational documentation current - failure modes and recovery procedures included.
- Own the operational health of our data stores - MongoDB Atlas backups and recovery, performance tuning, and monitoring at scale - so data stays durable, performant, and recoverable.
HOW YOU'L
You found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →