Weekend Site Reliability Engineer
Sporty Group
| Company | Sporty Group |
| Category | Uncategorised |
| Location | Global - Remote |
| Remote | Remote |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 6 Jan 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (greenhouse) |
Description
What you’ll be doing
Work with a team of DevOps and DBA professionals; covering Saturday, Sunday and Monday (5 days in total with flexibility in your days off) as a Weekend SRE
Improve existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in the future
Continuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practices
Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM)
Own weekend on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviews
Design and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification flooding
Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisation
Take ownership and responsibility for our cloud operation activities
Liaise with external security agencies for annual audits as well as perform our own internal security sweeps
Aid in reconfiguring existing architecture to allow for rapid deployments to new countries
Mentoring less experienced team members
What you’ll bring
3+ years DevOps / platform engineering experience
Must be based in Europe or Asia or LatAM
Experience independently leading the planning and deployment of a project
Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production
Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valued
Experience with Infrastructure-as-Code, particularly Terraform
Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus
Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetry
Experience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plus
Proven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions
Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns
Experience defining SLIs and SLOs and using them to inform reliability work
Familiarity with service mesh concepts is a plus, as we are actively evaluating Cilium-based service mesh in non-production environments
Solid networking knowledge, especially the TCP / IP stack and HTTP protocol
Experience handling high HTTP request volumes and designing systems for high availability and high traffic environments
A strong understanding of cache, including CDN, HTTP cache, Redis / Memcached
Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous
Our stack
Languages: Java / Spring Boot, Node.js, Python, JavaScript
Database: Aurora MySQL & PostgreSQL, MongoDB, MySQL Community
Cache: ElastiCache, Redis, Valkey
Messaging: Apache RocketMQ, AutoMQ, Kafka
Networking & Proxy: Nginx, Kong, Cilium, eBPF
Orchestration & GitOps: Docker, Kubernetes (EKS), ArgoCD, Helm
Computing & Storage: AWS EC2, VPC, AWS Lambda, EBS, S3
CI/CD: Jenkins, GitHub Actions
Metrics: Prometheus, Mimir, Grafana, Alertmanager
Logs: Loki, Vector
Traces: Tempo, OpenTelemetry, Alloy
Profiling: Pyroscope
RUM: Grafana
986,449 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →