Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Senior Platform Engineer (Reliability) - Unannounced Project

Scopely
CompanyScopely
CategoryEngineering
LocationES - Spain; GB - United Kingdom; IE - Ireland; PT - Portugal
RemoteOn-site (inferred)
EmploymentNot stated
LevelSenior
SalaryNot stated by the employer
Posted28 Jan 2026
Last verified2 Aug 2026
SourceEmployer ATS (greenhouse)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Scopely is looking for a  Senior Platform Engineer (Reliability) to join a new truly unique multiplayer strategy game in Spain, Ireland, Portugal or the UK on a remote or hybrid basis. We can support with visa sponsorship and relocation assistance from any location. At Scopely, we care deeply about what we do and want to inspire play every day — whether in our work environments alongside our talented colleagues or through our deep connections with our communities of players. We are a global team of game lovers who are developing, publishing and innovating the mobile games industry, connecting millions of people around the world daily. We are in the early stages of development on an ambitious, unannounced Strategy/MMO title, creating a team of talented and passionate game makers to join us on this exciting journey! What You'll Do Messaging Systems Ownership: Operate, monitor, and continuously improve our messaging infrastructure — with a current focus on NATS cluster and NATS JetStream — ensuring it is reliable, observable, and well-understood by the teams that depend on it. Signals & Observability: Design and own the observability layer for distributed backend systems, defining the signals (metrics, traces, logs) that make operational problems visible and actionable — and pushing for their adoption where they don't yet exist. Cross-functional Diagnosis: Sit at the intersection of infrastructure and backend engineering: correlate infrastructure signals (IOPS, latency, resource saturation) with application behaviour (message throughput, consumer lag, retry storms) to diagnose root causes and guide the right teams toward the right fixes. SLOs & Error Budgets: Define, implement, and maintain SLO related frameworks for backend services and messaging pipelines, making reliability measurable and helping teams make informed trade-offs between velocity and stability. Reliability as Internal Product: Build and maintain reliability tooling, runbooks, and operational frameworks as internal products — enabling backend and infrastructure engineers to self-serve on operational concerns rather than creating dependency on SRE. Partner with backend, infrastructure, and product engineers to shape reliability standards, share operational context, and influence architecture decisions before they become production problems. Incident Management: Lead or contribute to incident response across the messaging and backend layers, drive postmortems to systemic fixes, and embed preventative improvements into engineering workflows. Code Literacy & Engineering Collaboration: Navigate the infrastructure and backend codebases confidently; identify poorly instrumented services, inadequate infrastructure architecture, missing error handling, or patterns that create operational risk, contributing and improving  in close collaboration with other engineering teams. What We're Looking For Strong background in Site Reliability Engineering, production operations, or backend engineering with a significant operational focus Hands-on experience operating NATS and/or NATS JetStream in production — or equivalent deep experience with distributed messaging systems such as Apache Kafka, AWS Kinesis, or similar. Regardless of the system, experience with Leader Election, RAFT consensus, log replication, are essential. Ability to navigate and reason about application codebases (e.g. C#, Go, Python, or similar) while being capable of identifying instrumentation gaps, operational anti-patterns, and code-level root causes Strong observability experience: designing and implementing metrics, logs, and traces strategies across distributed systems Experience debugging complex distributed systems, particularly across the boundary between infrastructure and application layers Solid understanding of cloud infrastructure (AWS preferred) and containerized workloads (ECS, Kubernetes/EKS, or equival
HOUSE AD2,146,529 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →