Senior Platform Engineer (Reliability) - Unannounced Project
Scopely
| Company | Scopely |
| Category | Engineering |
| Location | ES - Spain; GB - United Kingdom; IE - Ireland; PT - Portugal |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 28 Jan 2026 |
| Last verified | 2 Aug 2026 |
| Source | Employer ATS (greenhouse) |
Description
Scopely is looking for a Senior Platform Engineer (Reliability) to join a new truly unique multiplayer strategy game in Spain, Ireland, Portugal or the UK on a remote or hybrid basis. We can support with visa sponsorship and relocation assistance from any location.
At Scopely, we care deeply about what we do and want to inspire play every day — whether in our work environments alongside our talented colleagues or through our deep connections with our communities of players. We are a global team of game lovers who are developing, publishing and innovating the mobile games industry, connecting millions of people around the world daily.
We are in the early stages of development on an ambitious, unannounced Strategy/MMO title, creating a team of talented and passionate game makers to join us on this exciting journey!
What You'll Do
Messaging Systems Ownership: Operate, monitor, and continuously improve our messaging infrastructure — with a current focus on NATS cluster and NATS JetStream — ensuring it is reliable, observable, and well-understood by the teams that depend on it.
Signals & Observability: Design and own the observability layer for distributed backend systems, defining the signals (metrics, traces, logs) that make operational problems visible and actionable — and pushing for their adoption where they don't yet exist.
Cross-functional Diagnosis: Sit at the intersection of infrastructure and backend engineering: correlate infrastructure signals (IOPS, latency, resource saturation) with application behaviour (message throughput, consumer lag, retry storms) to diagnose root causes and guide the right teams toward the right fixes.
SLOs & Error Budgets: Define, implement, and maintain SLO related frameworks for backend services and messaging pipelines, making reliability measurable and helping teams make informed trade-offs between velocity and stability.
Reliability as Internal Product: Build and maintain reliability tooling, runbooks, and operational frameworks as internal products — enabling backend and infrastructure engineers to self-serve on operational concerns rather than creating dependency on SRE. Partner with backend, infrastructure, and product engineers to shape reliability standards, share operational context, and influence architecture decisions before they become production problems.
Incident Management: Lead or contribute to incident response across the messaging and backend layers, drive postmortems to systemic fixes, and embed preventative improvements into engineering workflows.
Code Literacy & Engineering Collaboration: Navigate the infrastructure and backend codebases confidently; identify poorly instrumented services, inadequate infrastructure architecture, missing error handling, or patterns that create operational risk, contributing and improving in close collaboration with other engineering teams.
What We're Looking For
Strong background in Site Reliability Engineering, production operations, or backend engineering with a significant operational focus
Hands-on experience operating NATS and/or NATS JetStream in production — or equivalent deep experience with distributed messaging systems such as Apache Kafka, AWS Kinesis, or similar. Regardless of the system, experience with Leader Election, RAFT consensus, log replication, are essential.
Ability to navigate and reason about application codebases (e.g. C#, Go, Python, or similar) while being capable of identifying instrumentation gaps, operational anti-patterns, and code-level root causes
Strong observability experience: designing and implementing metrics, logs, and traces strategies across distributed systems
Experience debugging complex distributed systems, particularly across the boundary between infrastructure and application layers
Solid understanding of cloud infrastructure (AWS preferred) and containerized workloads (ECS, Kubernetes/EKS, or equival
2,146,529 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →