Senior Site Reliability Engineer
AlayaCare
| Company | AlayaCare |
| Category | Engineering |
| Location | Sydney |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 9 Jun 2026 |
| Last verified | 2 Aug 2026 |
| Source | Employer career page (greenhouse) |
Description
We’re hiring a Senior Site Reliability Engineer at AlayaCare
🗓️ Full-time | Permanent
📍 Preferred locations: Sydney, Brisbane, or Melbourne. Open to Perth based for the right candidate.
🏡 Hybrid working: 2 days in office, 3 days WFH
Competitive salary package with company stock, five wellness days per year, a flexible benefits package of $1000 per year, purposeful work.
👋 Meet AlayaCare! We’re a fast-growing SaaS scale-up on a mission to transform aged and disability care across Australia, Canada, the US and beyond. Our platform helps care providers deliver exceptional service in homes, communities, and residential settings. We’re big on Tech with Purpose and passionate about improving lives while having fun along the way.
The Role:
We’re on the lookout for a Senior Site Reliability Engineer who’s ready to bring their self-starting nature, AWS experience & analytical mind to the table. Reporting to the SRE, Engineering Manager, you'll help drive the reliability of our live SaaS solutions across the region.
Your days will involve:
Development, Automation, and Tooling
Design, build, and maintain infrastructure and platform services, including Kubernetes and observability tooling
Implement infrastructure as code, configuration management, and automated testing to ensure reliable, repeatable environments
Contribute to code and configuration reviews to improve scalability, maintainability, and reuse.
All the above using AI first mindset and development tooling such as Cursor and Kiro.
Build and tune AI Agents to accelerate delivery, and automate repetitive tasks.
Reliability and Operations
Monitor production systems, troubleshoot issues, and improve logging, monitoring, alerting, and runbooks
Participate in on-call rotations, incident response, and post-incident reviews to improve long-term reliability.
Requirements and Collaboration
Partner with Product, Engineering, and development teams to translate requirements into practical infrastructure solutions.
Identify risks related to operability, security, performance, and cost, and recommend appropriate trade-offs.
Continuous Improvement
Contribute to operational quality through runbooks, security practices, performance tuning, and process improvements
Proactively identify issues, raise concerns, and stay current with emerging SRE practices and technologies.
You’ll thrive in this role if you:
Bring 5+ years of experience in SRE/DevOps or a similar role
Believe that AI agents can be better than humans at certain tasks and are interested in maximizing their usage
Have solid hands-on experience with AWS and Terraform
Have practical experience running workloads on Docker & Kubernetes
Are known for your problem-solving skills and your ability to work well autonomously
Are proficient in at least one development or scripting language (such as Python, Go, Bash)
Have practical knowledge of infrastructure as a code (CloudFormation or Terraform)
Have some knowledge of APM, logging & metrics systems (New Relic, Prometheus or ELK)
Have some background knowledge of system & network security fundamen
1,868,060 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →