Senior SRE Engineer
Codeway
| Company | Codeway |
| Category | Engineering |
| Location | Barcelona |
| Remote | Hybrid |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 15 Jul 2026 |
| Last verified | 3 Aug 2026 |
| Source | Employer career page (ashby) |
Description
ABOUT CODEWAY
Codeway is a global consumer tech company with more than 400M users worldwide.
Since 2020, we’ve built and scaled 60+ mobile apps across creativity, productivity, wellness, language learning, and entertainment.
Our flagship apps — Retake AI, Cleanup, Learna, and DramaPops — and many of them lead their categories globally. In 2024, we became the most downloaded app publisher on iOS, driven by cutting-edge AI research, sharp data-driven execution, and a relentless focus on product and marketing.
We’re a team of 300+ people across İstanbul and Barcelona who bring curiosity, passion, trust, and ownership to everything we build. Recognized as a #1 LinkedIn Top Startup and a Great Place to Work in Europe, Codeway is where ambitious people do their life’s best work.
We’re building the next generation of consumer tech and reimagining what mobile apps can be.
This is Codeway. This is our way. Join us.
POSITION
We’re looking for a Senior Site Reliability Engineer to own and mature reliability, performance, and security across our growing platform. This role sits at the intersection of Engineering, Infrastructure, and Security, helping design, operate, and continuously improve the systems that keep dozens of consumer apps running for users around the world.
You’ll work closely with product engineering teams to make reliability measurable rather than assumed. That means defining and enforcing SLIs, SLOs, and error budgets; operating and hardening a multi-cluster Kubernetes environment; building the observability that catches problems before users feel them; and leading incident response when things break. It’s a hands-on role with real ownership over how reliability and security evolve as we scale.
Several parts of our reliability practice are still early in their maturity. We’re looking for someone who enjoys building the standards, processes, tooling, and automation that will form the foundation of our SRE function — not someone waiting for a playbook to already exist.
We welcome applicants from all backgrounds and experiences. If you’re excited about running systems at consumer scale and believe you could be a strong fit, we encourage you to apply, even if your experience doesn’t align perfectly with every qualification listed below.
WHAT YOU’LL BE DOING
Reliability, SLOs, Observability
- Define, instrument, and report on SLIs, SLOs, and error budgets across critical services, so reliability decisions are driven by data rather than opinion.
- Own observability end-to-end — metrics, logs, traces, dashboards, and alerting — and drive measurable reductions in detection and resolution times.
- Reduce alert noise and false positives so on-call engineers can trust what wakes them up.
- Run reliability reviews and an error-budget policy that shapes how teams prioritize between shipping and stability.
Kubernetes & Platform Operations
- Operate, scale, and upgrade our multi-cluster Kubernetes (GKE) environment: cluster lifecycle, autoscaling, networking, ingress, and resource management.
- Act as the deep-expertise escalation point for cluster and platform issues across dozens of services.
- Own capacity planning, performance, and cloud cost efficiency, balancing spend against reliability targets.
- Build self-service platform tooling that lets product teams move quickly without needing to become infrastructure experts.
Security & Resilience
- Embed security into the platform through RBAC and least-privilege, secrets management, image and dependency scanning, network policies, and a disciplined patching cadence.
- Partner with the security function on vulnerability remediation, audit readiness, and secure-by-default infrastructure.
- Own disaster recovery: define and regularly validate RTO/RPO targets through DR drills and failure testing.
- Contribute to architecture and production-readiness reviews so reliability and security are designed in, not bolted o