Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Senior Platform Engineer

Element451
CompanyElement451
CategoryEngineering
LocationRemote
RemoteRemote
EmploymentNot stated
LevelSenior
SalaryNot stated by the employer
Posted2 Jul 2026
Last verified30 Jul 2026
SourceEmployer career page (ashby)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Element451 is building the AI-powered platform reshaping how colleges and universities recruit, enroll, and support their students - and the reliability of that platform is what earns the trust of the institutions that run on it. This is a rare chance to own reliability, operations, and security at the level where it actually happens: keeping a real production system healthy, fast, and safe, and building the automation that keeps it that way. If you're a broad operator who's happiest when a single week spans operations, delivery, security, infrastructure, and data - and you've done that work where things move fast and the edges go unowned - you'll feel at home here. THE ROLE You'll keep Element451's platform reliable, secure, and operable - and build the delivery systems that let it scale without scaling the firefighting. This is a hands-on senior IC role, and a deliberately broad one: reliability and operations are the core, with CI/CD and delivery, security, infrastructure, and data reliability all real, recurring parts of the work. You'll partner closely with our Director of Platform Engineering, who owns the platform strategy; you own a large share of the operational and delivery execution that brings it to life. We hold a high bar - and we give you the ownership, context, and support to meet it. The work has range and a fair amount of unpredictability; the people who thrive in it tend to want exactly that. WHAT YOU'LL OWN You own the operational health of the platform - that it's available, fast, observable, and safe in production - and you build the automation that makes those properties durable rather than heroic. In practice, you: - Own the reliability discipline in practice - define and track SLIs and SLOs, keep the observability stack sharp (we use CloudWatch, Sentry, Papertrail, and Langsmith), and make system health and customer impact legible in real time. - Carry production operations day to day - participate in on-call, lead incident response, run blameless post-incident reviews, and drive issues to root cause rather than patching symptoms. - Treat operational toil as engineering work to eliminate - relentlessly automate remediation, sharpen alert quality, and drive down MTTD and MTTR rather than absorbing manual load. - Own and evolve the CI/CD and delivery platform alongside the Director of Platform Engineering - build pipelines, deployment automation, environment management, and release tooling - so shipping is routine, safe, and low-drama, with progressive delivery, automated rollback, and production validation gates as standard. - Build the developer-facing automation and paved roads that cut friction for the product engineering team - treating them as the platform's customer and making the reliable, secure path the easy path. - Be the platform's hands-on security operator - IAM and least-privilege hygiene, secrets management, threat detection and response (WAF, GuardDuty), and vulnerability triage and remediation against SLA. This is the security function in practice today; you partner with the Director of Platform Engineering on strategy and standards, and you're trusted to set the operational bar where none exists yet. - Keep the platform audit-ready by default - produce the infrastructure and reliability evidence that SOC 2 Type II and FERPA obligations depend on (we use Vanta), so audits are a byproduct of good operations, not a scramble. - Build and operate cloud infrastructure as code - AWS (ECS/Fargate, Lambda, SQS/SNS, EventBridge, S3/CloudFront, VPC) managed with Terraform, with no manual or snowflake infrastructure in any environment. - Plan and execute scaling ahead of product growth, and keep operational documentation current - failure modes and recovery procedures included. - Own the operational health of our data stores - MongoDB Atlas backups and recovery, performance tuning, and monitoring at scale - so data stays durable, performant, and recoverable. HOW YOU'L
HOUSE ADYou found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →
Senior Platform Engineer — Element451 · Job Opportunities API