Site Reliability Engineer
Man Group
| Company | Man Group |
| Category | Engineering |
| Location | Sofia |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 5 Dec 2025 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (greenhouse) |
Description
About Man Group
Man Group is a global alternative investment management firm focused on pursuing outperformance for sophisticated clients via our Systematic, Discretionary and Solutions offerings. Powered by talent and advanced technology, our single and multi-manager investment strategies are underpinned by deep research and span public and private markets, across all major asset classes, with a significant focus on alternatives. Man Group takes a partnership approach to working with clients, establishing deep connections and creating tailored solutions to meet their investment goals and those of the millions of retirees and savers they represent.
Headquartered in London, we manage $228.7 billion* and operate across multiple offices globally. Man Group plc is listed on the London Stock Exchange under the ticker EMG.LN and is a constituent of the FTSE 250 Index. Further information can be found at www.man.com
At Man Group, we respect your privacy and we are committed to protecting and safeguarding your Personal Data. We have developed policies and processes which are designed to provide for the security and integrity of your Personal Data. We are committed to Processing your Personal Data fairly and lawfully, and being open and transparent about such Processing. For further information on how we process your data, please see the privacy notice for applicants here
* As at 31 March 2026
The Role
Join our high-performing Site Reliability Engineering (SRE) team and play a pivotal role in ensuring the reliability, scalability, and performance of the technology powering Man Group’s hedge funds. You’ll have the autonomy, tools, and support to innovate and shape the future of our platform. This is an opportunity to work on cutting-edge projects, gain mentorship from senior leaders, and develop a deep understanding of both technology and the business.
As an SRE, you’ll take ownership of service reliability and deliver solutions that make a real impact. Your initial focus will include leveraging AI to accelerate incident diagnosis and resolution, improving observability, capacity planning, and automation. Over time, you’ll work across our entire infrastructure stack, operating at scale and driving continuous improvement.
Role Responsibilities
Ensure reliability and performance of critical systems across global infrastructure through proactive monitoring and rapid incident response.
Design and implement observability solutions using tools like Prometheus, Grafana, ELK, and Loki to provide deep insights into system health.
Automate operational tasks and build self-service capabilities to eliminate toil and improve efficiency.
Develop and maintain SLIs, SLOs, and error budgets to guide reliability improvements and inform engineering priorities.
Participate in incident response efforts, blameless post-mortems, and implement preventive measures to reduce recurrence.
Collaborate with development teams to improve system design, deployment practices, and operational excellence.
Operate at scale, managing petabyte-level storage, large CPU/GPU deployments, and high-throughput distributed systems.
Contribute to capacity planning and performance tuning, ensuring systems meet business demands.
Manage multiple ELK clusters hosting hundreds of terabytes of logs, telemetry, and APM data.
Key competencies
Required
Strong understanding of SRE principles, including SLIs, SLOs, error budgets, and reliability best practices.
Hands-on experience with observability and monitoring tools (Prometheus, Grafana, ELK, Loki, or similar).
Proficiency with automation tools (Ansible, Terraform) and scripting/programming languages (Python, Go, PowerShell).
Strong troubleshooting and debugging skills across distributed systems, with the ability to diagnose complex production issues under pressure.
Experience with incident management, on-call rotations, and post-incident reviews
You found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →