Senior Site Reliability Engineer
ZigZag Careers
| Company | ZigZag Careers |
| Category | Engineering |
| Location | — |
| Remote | Remote |
| Employment | Contract |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 15 Jul 2026 |
| Last verified | 12 Aug 2026 |
| Source | The employer's own careers page (company_site) |
Description
Remote, Philippines. Full-time.OverviewAs a Senior Site Reliability Engineer, you will design, build, and maintain the infrastructure and automation that power the platform. Working closely with software engineering teams and SRE peers, you will embed reliability, performance, and compliance into the development lifecycle, with focus on scalability, resilience, security, and operational efficiency across all environments.Key ResponsibilitiesReliability Engineering & Operational ExcellenceDesign, implement, and continuously improve highly available, scalable, secure, and resilient cloud infrastructure and platform services.Define and evolve SLIs, SLOs, and operational metrics to drive measurable reliability outcomes.Lead incident response, major incident management, root cause analysis, and post-incident reviews focused on systemic improvement.Drive reduction of operational toil through automation, standardisation, and self-healing platform capabilities.Develop and maintain disaster recovery, backup, failover, and resilience strategies to meet defined RTO and RPO objectives.Conduct capacity planning, performance analysis, and proactive optimisation of infrastructure and application environments.Platform & Infrastructure EngineeringArchitect, build, and maintain scalable cloud-native infrastructure primarily within AWS.Develop and maintain infrastructure-as-code using Terraform and CloudFormation.Build reusable platform components and shared services that improve developer productivity.Develop automation tooling and operational frameworks using Python.Ensure infrastructure configurations, architecture decisions, and operational processes are thoroughly documented and auditable.Observability, Monitoring & PerformanceDesign and maintain comprehensive observability solutions covering metrics, logging, tracing, alerting, and dashboarding.Improve platform visibility using tools such as AWS CloudWatch, Sumo Logic, Datadog, or Grafana.Develop actionable alerting strategies that reduce noise and improve incident response effectiveness.Drive adoption of observability best practices across engineering teams.CI/CD & Developer EnablementDesign and enhance robust CI/CD pipelines and deployment strategies supporting safe, reliable software delivery.Support progressive delivery practices including blue/green deployments, canary releases, and zero-downtime deployments.Collaborate with engineering teams to embed reliability, scalability, and security into the SDLC.Security, Risk & ComplianceSupport vulnerability management and remediation using tools such as Snyk, Lacework, or Tenable Nessus.Assist in maintaining compliance with PCI-DSS, ISO27001, SOC 2, and internal security controls.Contribute to security hardening, access management, and audit readiness initiatives.Leadership & CollaborationAct as a technical leader and mentor within the SRE and broader engineering teams.Contribute to engineering standards, operational best practices, and platform strategy.Collaborate with cross-functional stakeholders including Engineering, Product, Security, Architecture, and external vendors.Skills & Experience5+ years in Site Reliability Engineering, DevOps Engineering, Platform Engineering, or related infrastructure roles.Strong hands-on experience with production workloads in AWS.Deep experience with Terraform and/or CloudFormation.Strong experience designing and supporting CI/CD pipelines.Strong understanding of distributed systems, microservices architecture, networking, and cloud-native technologies.Strong scripting and automation skills in Python, Bash, or similar.Experience managing production incidents and conducting structured root cause analysis.DesirableExperience with Kubernetes, container orchestration, and platform engineering.Experience with Buildkite, GitHub Actions, or GitLab CI.Exposure to service mesh, event-driven architectures, and distributed tracing.Experience with PCI-DSS, ISO27001, or SOC 2 compliance frameworks.Experience with FinOps, cloud cost optimisation, and infrastructure performance tuning.Familiarity with DevSecOps practices.Experience mentoring engineers or leading technical initiatives.