Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Senior Site Reliability Engineer (SRE)

Taboola
CompanyTaboola
CategoryEngineering
LocationTel Aviv
RemoteOn-site (inferred)
EmploymentNot stated
LevelSenior
SalaryNot stated by the employer
Posted26 Jul 2026
Last verified12 Aug 2026
SourceThe employer's own careers page (company_site)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Taboola is a leading performance advertising company seeking a Senior Site Reliability Engineer to join the R&D Infrastructure team in Tel Aviv. You will build, scale, and maintain high-scale hybrid infrastructure spanning on-premise, public cloud, and AI/ML Kubernetes environments, focusing on reliability, performance, and cost efficiency. What You'll Do • Maintain hybrid infrastructure (on-prem, public cloud, AI/ML clusters) with high availability, performance, and cost efficiency • Build internal software tooling and manage Infrastructure as Code pipelines using Go, Python, or Rust to eliminate repetitive operations • Perform deep-dive troubleshooting across the full stack from CDN edge configurations to Linux kernel tuning and network layer bottlenecks • Design and maintain monitoring and alerting setups to proactively identify and address system health issues • Participate in on-call rotations, lead incident resolution, and conduct blameless post-mortems to ensure system resilience What You Need • 7+ years of experience managing, scaling, and troubleshooting large-scale distributed Linux environments in production • Deep understanding of Linux system internals and network protocols (TCP/IP, DNS, HTTP, gRPC) • Hands-on experience with edge/CDN services such as Fastly, Cloudflare, Akamai, or CloudFront • Hands-on experience with Infrastructure as Code and orchestration tools such as Terraform, Ansible, Puppet, ArgoCD, or Jenkins • Production experience managing containerized environments using Kubernetes and Docker • Solid programming skills in at least one modern language (Go, Python, or Rust) Nice to Have • Experience designing and operating telemetry, metrics collection, and alerting stacks at scale (Prometheus, Grafana, ELK/logging) • Practical background in optimizing infrastructure costs and resource efficiency across cloud and on-prem Comprehensive benefits including health coverage, fully stocked kitchen, and location-specific perks (gym partnerships, parking). Hybrid work schedule with 3 days in-office.