Senior Site Reliability Engineer (SRE)
Taboola
| Company | Taboola |
| Category | Engineering |
| Location | Tel Aviv |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 26 Jul 2026 |
| Last verified | 12 Aug 2026 |
| Source | The employer's own careers page (company_site) |
Description
Taboola is a leading performance advertising company seeking a Senior Site Reliability Engineer to join the R&D Infrastructure team in Tel Aviv. You will build, scale, and maintain high-scale hybrid infrastructure spanning on-premise, public cloud, and AI/ML Kubernetes environments, focusing on reliability, performance, and cost efficiency.
What You'll Do
• Maintain hybrid infrastructure (on-prem, public cloud, AI/ML clusters) with high availability, performance, and cost efficiency
• Build internal software tooling and manage Infrastructure as Code pipelines using Go, Python, or Rust to eliminate repetitive operations
• Perform deep-dive troubleshooting across the full stack from CDN edge configurations to Linux kernel tuning and network layer bottlenecks
• Design and maintain monitoring and alerting setups to proactively identify and address system health issues
• Participate in on-call rotations, lead incident resolution, and conduct blameless post-mortems to ensure system resilience
What You Need
• 7+ years of experience managing, scaling, and troubleshooting large-scale distributed Linux environments in production
• Deep understanding of Linux system internals and network protocols (TCP/IP, DNS, HTTP, gRPC)
• Hands-on experience with edge/CDN services such as Fastly, Cloudflare, Akamai, or CloudFront
• Hands-on experience with Infrastructure as Code and orchestration tools such as Terraform, Ansible, Puppet, ArgoCD, or Jenkins
• Production experience managing containerized environments using Kubernetes and Docker
• Solid programming skills in at least one modern language (Go, Python, or Rust)
Nice to Have
• Experience designing and operating telemetry, metrics collection, and alerting stacks at scale (Prometheus, Grafana, ELK/logging)
• Practical background in optimizing infrastructure costs and resource efficiency across cloud and on-prem
Comprehensive benefits including health coverage, fully stocked kitchen, and location-specific perks (gym partnerships, parking). Hybrid work schedule with 3 days in-office.