Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Senior Cloud Site Reliability Engineer

NICE
CompanyNICE
CategoryEngineering
LocationPune
RemoteOn-site (inferred)
EmploymentNot stated
LevelSenior
SalaryNot stated by the employer
Posted29 Jul 2026
Last verified12 Aug 2026
SourceThe employer's own careers page (company_site)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
NICE is a global software leader seeking a Senior Site Reliability Engineer to join its core Reliability Engineering team. This role focuses on ensuring scalability, reliability, and performance of mission-critical systems and observability platforms across multiple cloud environments and regions, with strong emphasis on automation, cloud-native operations, and incident management. What You'll Do • Design and implement scalable, reliable, and resilient systems across hybrid or multi-cloud environments (AWS/EKS/ECS/Lambda) • Build and manage infrastructure automation using Terraform, Helm, and Kubernetes; improve CI/CD pipelines with Jenkins and GitHub Actions • Own and enhance the observability stack (Prometheus, Grafana, Loki, Tempo, OpenTelemetry, Mimir) and define SLOs and error budgets • Lead major incident response, root cause analysis, and blameless postmortems; partner with product teams on operational readiness • Mentor junior SREs and developers on reliability practices, automation, and observability; contribute to technical roadmaps and reliability-focused design reviews What You Need • 5+ years of strong experience with Kubernetes, EKS, ECS, and containerized workloads in production • Expertise in AWS services (EC2, Lambda, IAM, RDS, S3, ALB/NLB, VPC, PrivateLink) • Proficiency with Terraform, Helm, Jenkins, GitHub Actions, and GitOps tools (ArgoCD or Flux) • Deep understanding of observability frameworks (metrics, logs, traces, distributed monitoring) and hands-on experience with Prometheus, Grafana, Loki, Tempo, Alloy, and OpenTelemetry • Strong knowledge of Linux, networking fundamentals, system performance tuning, and proficiency in Python, Go, or Shell scripting for automation