Principal Site Reliability Engineer, Machine Learning
Cambridge Mobile Telematics
| Company | Cambridge Mobile Telematics |
| Category | Uncategorised |
| Location | Cambridge |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 24 Jul 2026 |
| Last verified | 2 Aug 2026 |
| Source | Employer ATS (greenhouse) |
Description
CMT is looking for a Principal Site Reliability Engineer I, Machine Learning to help us change the world. CMT has helped protect over 65 million drivers and prevent over 126,000 crashes worldwide. We build AI to solve some of the most difficult challenges in mobility — understanding and reducing risk, detecting crashes, and getting people life-saving help. The problems are hard. The impact is real. No matter your role, your work will matter at CMT.
CMT is looking for a collaborative, customer-committed, and creative SRE team member who wants to join us in making roads safer by making drivers better!
Responsibilities:
Use independent judgment and discretion to own SLOs, error budgets, and the operational health of Ray clusters running on AWS EKS and Databricks workloads on AWS EC2 across multiple accounts and regions
Maintain the observability of uptime, availability, and scalability of EKS Ray and Databricks workloads using CloudWatch and Datadog, including defining alerting that maps to SLOs
Operate and tune EKS Ray workloads at scale including autoscaling, GPU scheduling, and automated failure recovery
Manage Databricks on AWS including workspace administration, cluster policies, Unity Catalog, job orchestration, and IAM Roles and Policies
Maintain ongoing cost visibility, cost optimization, and capacity planning across EC2 and EKS workloads, including through the use of On Demand Capacity Reservations and Spot lifecycle
Perform ongoing maintenance of the underlying EC2 and EKS infrastructure, including regular security updates and operating system upgrades
Codify everything as infrastructure-as-code using Terraform and CI/CD pipelines, enabling updates through Pull Requests with approval workflows, while also automating maintenance tasks to reduce toil
Lead incident response for Data Science and Machine Learning platform outages, run blameless postmortems, and drive systemic remediation, including participating in an on-call rotation
Complete any additional tasks as they arise
Qualifications:
Bachelor’s degree or equivalent years of experience and/or certification in a related field
7+ years working in Site Reliability Engineering or Information Technology
Design and document systems, including writing and reviewing code, to automate away problems within your team’s domain
Intermediate to expert experience deploying and maintaining AWS services such as EC2, ECS, EKS, SQS, Lambda, Dynamo, RDS/Aurora, S3, and IAM
Intermediate to expert experience monitoring services and applications using tools such as CloudWatch Metrics, CloudWatch Logs, and Datadog, including defining and configuring alerts and SLO reports
Intermediate to expert experience maintaining the uptime and scalability of AWS compute services used for Machine Learning and Data Science workloads, specifically EC2 and EKS
Intermediate to expert coding skills in at least one programming language; we work primarily in Python
Intermediate to expert experience using Infrastructure as Code platforms and CI/CD pipelines, specifically Terraform, to manage AWS infrastructure and services
In-depth knowledge & experience with Linux operating systems (Amazon Linux, Ubuntu) on EC2 and Docker / Kubernetes
Experience with leading projects in system design, architecture changes, and technology selection
Compensation and Benefits:
Fair and competitive salary based on skills and experience, and annual performance bonus
Equity may be awarded in the form of Restricted Stock Units (RSUs)
Medical, Dental, Vision and Life Insurance, matching 401k, short-term & long-term disability and parental leave
Unlimited Paid Time Off including vacation, sick days & public holidays
Flexible scheduling and work from home policy depending on role and responsibilities
Base Salary Range
The base salary range for this position is: $142,000