Site Reliability Engineer I
Zafin
| Company | Zafin |
| Category | Engineering |
| Location | India - Trivandrum |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 14 May 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (greenhouse) |
Description
Zafin is an AI platform company helping regulated institutions modernize how critical work is designed, governed, and delivered. Our technology enables organizations to move faster while maintaining the governance, accountability, and control required in highly regulated environments.
Our portfolio includes Zafin AIOS , an agent orchestration platform for governed AI work; the Zafin Banking Platform , which helps banks modernize product, pricing, offers, billing, loyalty, and relationship management; and Zafin IO , an integration platform that connects data, systems, and workflows across complex enterprise environments.
Headquartered in Toronto, Canada, Zafin partners with leading financial institutions across North America, Europe, the Middle East, Africa, and Asia-Pacific. As AI transforms the future of financial services, we're building the platforms that help regulated organizations adopt AI responsibly and at scale. Cloud Site Reliability Engineer I (CSRE I)
Zafin is seeking a Cloud Site Reliability Engineer I (CSRE I) to lead strategic initiatives in ensuring the reliability, scalability, and performance of our cloud infrastructure and applications. This advanced role requires mastery in cloud technologies, strategic planning, and incident management to drive innovative solutions and operational excellence.
Key Responsibilities
Manage the resolution of complex technical issues involving Zafin’s products and Azure cloud environment.
Design and implement strategic operational enhancements to improve resiliency and system reliability.
Conduct in-depth Root Cause Analysis (RCA) for high-severity incidents and drive initiatives to reduce error recurrence.
Represent the organization in external client escalation calls, providing expert guidance and solutions.
Optimize cloud infrastructure for high performance, scalability, and cost-effectiveness.
Provide thought leadership in managing and scaling container orchestration platforms such as AKS and OpenShift.
Oversee the implementation of advanced monitoring solutions and integrate predictive analytics for proactive issue resolution.
Develop and execute automation strategies to streamline operational workflows and incident responses.
Create and maintain comprehensive documentation of cloud architectures, processes, and incident management strategies.
Mentor and coach junior engineers, fostering a culture of continuous learning and innovation.
Drive strategic initiatives, collaborating with cross-functional teams to achieve organizational objectives.
Qualifications
Bachelor’s degree in computer science, Engineering, or a related field (Master’s degree preferred).
8+ years of experience in cloud support, operations, or a related role.
Advanced expertise in Microsoft Azure (preferred) or equivalent cloud platforms.
Demonstrated experience in designing and scaling container orchestration systems like AKS or OpenShift.
Proven leadership in managing automated deployment pipelines, including Azure DevOps.
Mastery in enterprise monitoring platforms (e.g., Azure Insights, Grafana) and predictive analytics tools.
Advanced scripting skills with PowerShell, Python, or similar languages.
Extensive experience in incident management and defining SLAs for global production environments.
In-depth knowledge of database management, particularly Postgres.
Preferred Qualifications
Advanced certifications in cloud platforms (e.g., Azure Solutions Architect Expert).
Experience with ITSM tools and proce
You found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →