Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Site Reliability Engineer

Allwyn UK
CompanyAllwyn UK
CategoryEngineering
LocationWatford
RemoteOn-site (inferred)
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted28 Jul 2026
Last verified1 Aug 2026
SourceEmployer career page (workable)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
At the heart of everything we do is our vision to change lives every day, and our mission to grow The National Lottery responsibly and champion its impact.     We are Allwyn UK, part of the Allwyn Entertainment Group – a multi-national lottery operator with a market-leading presence across the USA (Michigan and Illinois) and Europe, including Czech Republic, Austria, Greece, Cyprus and Italy.    While the main contribution of The National Lottery to society is through the funds to good causes, at Allwyn we put our purpose and values at the heart of everything we do.  Join us as we embark on a once-in-a-lifetime, largescale transformation journey by creating a National Lottery that delivers more money to good causes.      We’ll talk a bit more about us further down the page, but for now – let’s talk about the role and who we’re looking for…    A bit about the role  At Allwyn, the Site Reliability Engineer supports the reliability and performance of digital services by operating production systems, building automation, and improving observability. You will work closely with Senior SREs and engineering teams to ensure services remain stable, scalable, and well-instrumented across both steady-state and high-demand events. What you’ll be doing  Objectives of the role Maintain reliable production services across digital platforms Improve monitoring, alerting, and observability coverage Reduce operational toil through automation Support incident response and continuous improvement Contribute to performance and scaling of services Production operations Participate in 1-in-4 on-call rotation: Respond to incidents Support out-of-hours diagnosis Monitor system health across: Web and mobile platforms Games platforms Player services Troubleshoot issues across application, infrastructure, and network layers Incident response & improvement Support incident triage and resolution Participate in post-incident reviews and implement remediation actions Maintain and improve runbooks and operational documentation Observability Implement and maintain monitoring using: Splunk CloudWatch Grafana Improve: Logging quality Metrics coverage Alerting accuracy Contribute to linking system performance to user experience signals Automation & engineering Develop scripts and tooling to: Reduce manual tasks Improve repeatability and reliability Contribute to infrastructure management using Terraform Support deployment processes and CI/CD improvements Platform & cloud Work with AWS services (primarily ECS, with exposure to EKS/Kubernetes) Support scalability and availability improvements Assist in performance tuning and capacity planning Collaboration Work closely with engineers to: Improve service reliability Support releases and production readiness Contribute to adoption of SRE practices within teams What experience we’re looking for   Technical Experience in cloud environments (AWS preferred) Working knowledge of: Containers (ECS; exposure to Kubernetes is a plus) Infrastructure as Code (Terraform) Basic programming/scripting capability (Python, Bash, etc.) Troubleshooting & operations Ability to diagnose issues in distributed systems Familiarity with monitoring and logging tools Understanding of Linux systems and networking fundamentals SRE fundamentals Understanding of: Monitoring and alerting concepts Reliability principles Incident response processes Willingness to be part of an on-call rotation Desirable Experience Exposure to: SLOs / SLIs / error budgets CI/CD pipelines Performance and load testing Familiarity with observability platforms used in modern digital estates Experience
HOUSE AD1,229,276 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →
Site Reliability Engineer — Allwyn UK · Job Opportunities API