Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Senior Site Reliability Engineer

Aerospike
CompanyAerospike
CategoryEngineering
LocationAustralia
RemoteOn-site (inferred)
EmploymentNot stated
LevelSenior
SalaryNot stated by the employer
Posted19 Jun 2026
Last verified2 Aug 2026
SourceEmployer ATS (greenhouse)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Aerospike is the real-time database for mission-critical use cases and workloads, including machine learning, generative, and agentic AI. Aerospike powers millions of transactions per second with millisecond latency, at a fraction of the total cost of ownership compared to other databases. Global leaders, including Adobe, Airtel, Barclays, Criteo, DBS Bank, Experian, Grab, HDFC Bank, PayPal, Sony Interactive Entertainment, The Trade Desk, and Wayfair, rely on Aerospike for customer 360, fraud detection, real-time bidding, profile stores, recommendation engines, and other use cases.   At Aerospike, we dream big and deliver even bigger. Our mission is to unleash the power of the world’s real-time data with a database built for infinite scale, speed, and sustainability . If you're ready to shape the future of data, join us. Senior Site Reliability Engineer As a Senior Site Reliability Engineer (SRE) for Aerospike, you will be instrumental in designing, building, and optimizing a scalable, highly resilient cloud platform. You will focus on improving reliability, performance, and automation to ensure seamless delivery and operation of our cloud platform services. Your responsibilities will include developing robust infrastructure, implementing intelligent monitoring systems, and driving continuous improvement initiatives that enhance system efficiency, scalability, and overall platform stability. Key Responsibilities Designing, deploying, and optimizing large-scale Aerospike cloud platform infrastructure and services across multiple environments Leading the development and enhancement of automation and infrastructure-as-code solutions to improve operational efficiency Building and maintaining monitoring, alerting, and observability implementations to proactively detect and resolve system issues Leading incident response activities, conducting post-mortems, and driving continuous improvement initiatives Designing and enforcing security best practices for cloud infrastructure and access control Collaborating with development teams to ensure reliable service delivery and alignment with SRE best practices Participating in on-call rotation, responding to critical incidents and minimizing downtime through proactive mitigation strategies Establishing documentation standards, runbooks, and system configurations for team knowledge sharing Leading capacity planning and performance optimization efforts Mentoring junior engineers and sharing knowledge to build team capabilities Required Experience 6+ years of experience in Site Reliability Engineering (SRE), DevOps, or related fields, with a focus on building scalable, resilient, and automated cloud-based systems Hands-on experience designing, deploying, and optimizing production-grade, business-critical systems in cloud environments Expertise with at least one major public cloud provider (AWS, Google Cloud, or Azure), including cloud-native services and architectures Strong proficiency in infrastructure-as-code (IaC) tools such as Terraform to enable automated and reproducible infrastructure Experience in CI/CD pipeline design and implementation, enabling seamless, automated software delivery and infrastructure updates Deep understanding of Linux/Unix systems, networking fundamentals, and distributed system architectures Proficiency in scripting and software development using Python, Bash, or Go to build automation, tooling, and infrastructure enhancements Experience with containerization and orchestration technologies such as Docker and Kubernetes for efficient service deployment and scaling Hands-on experience with monitoring, logging, and observability tools (e.g., Prometheus, Grafana, Datadog, Elasticsearch, Kibana) to drive data-driven system improvements Strong problem-solving skills with an engineering-first mindset for improving system reliability, scalabi
HOUSE AD2,112,231 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →