Senior Site Reliability Specialist (SRE)
AlayaCare
| Company | AlayaCare |
| Category | Engineering |
| Location | Montréal |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 17 Jul 2026 |
| Last verified | 7 Aug 2026 |
| Source | Employer ATS (greenhouse) |
Description
About AlayaCare
At AlayaCare, we’re more than just a fast-growing SaaS company, we’re a team of people passionate about transforming home healthcare. Our cloud-based platform empowers care providers around the world to deliver better outcomes for their clients.
With 550+ employees across Canada, the US, Australia, and Brazil, we’re united by a shared mission and a strong culture of transparency, growth, and human connection. Whether you're early in your career or a seasoned expert, AlayaCare offers the opportunity to grow your impact, your skills, and your career.
About the Role
We are seeking a Senior Site Reliability Specialist to join our SRE team. Reporting to the Engineering Manager, you is responsible for scaling AWS cloud infrastructure, evolving Kubernetes deployment pipelines, improving monitoring, alerting, and resiliency, and developing tooling that enables product teams to deliver safely and efficiently. For acquired Azure-based products the focus is on monitoring, alert triage, and runbook-driven incident response rather than greenfield platform design.
This role owns shared platform services across cloud regions, including databases, messaging, logging, search, and tenant provisioning. The Senior SRE is expected to lead major infrastructure initiatives and proof-of-concept efforts, contribute to technical planning and prioritization, as well as partners with the Product teams to reduce operational incidents.
The SRE team also develops and operates AI-driven tools to streamline runbooks, accelerate incident response, and generate operational insights from platform telemetry to improve reliability and reduce manual effort.
What You’ll Do
Development, Automation, and Tooling
Design, build, and maintain infrastructure and platform services, including Kubernetes and observability tooling.
Implement infrastructure as code, configuration management, and automated testing to ensure reliable, repeatable environments.
Contribute to code and configuration reviews to improve scalability,maintainability, and reuse.
Reliability and Operations
Monitor production systems, troubleshoot issues, and improve logging, monitoring, alerting, and runbooks.
Participate in on-call and help desk rotations, incident response, and post-incident reviews to improve long-term reliability.
Requirements and Collaboration
Partner with Product, Engineering, and development teams to translate requirements into reliable and operable infrastructure solutions.
Identify risks across operability, security, performance, and cost, and recommend practical trade-offs.
Continuous Improvement
Contribute to operational quality through runbooks, security hardening, performance tuning, and process improvements.
Stay current with emerging SRE practices, including AI-assisted operations and modern AWS platform patterns.
What You Bring to the Team
Bachelor’s or advanced degree in computer science, computer engineering, or related practical fields with demonstrated experience.
5+ years of hands-on experience
Solid hands-on experience with AWS in a multi-account, multi-region environment: EKS, AWS Organizations, IAM, and KMS.
Strong proficiency with Terraform and Infrastructure as Code workflows, including Atlantis/GitOps, state management, and module/provider upgrades.
Practical experience running workloads on Docker and Kubernetes in production, including Gateway API ingress patterns and cluster lifecycle management (upgrades, addons, node provisioning).
Strong experience with Linux systems administration and production troubleshooting.
Proficiency in at least one development or scripting language such as Python, Go, or Bash.
Experience with an observability platform (e.g. New Relic, OpenSearch, CloudWatch, OpenTelemetry) and event-driven alerting (e.g. EventBridge, SNS, PagerDut