Platform Site Reliability Engineer
Radiant
| Company | Radiant |
| Category | Engineering |
| Location | Gloucestershire |
| Remote | Hybrid |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 6 May 2026 |
| Last verified | 3 Aug 2026 |
| Source | Employer ATS (ashby) |
Description
ABOUT RADIANT
Radiant is redefining how AI infrastructure is built.
We design and operate AI-native cloud platforms engineered for sovereignty, performance, and scale. Our infrastructure powers GPU-native workloads, multi-tenant control planes, and high-performance AI systems designed for the most demanding environments.
We are not building a generic cloud. We are building purpose-built AI infrastructure - from powered land, to compute, to software .
As we scale our platform and expand our engineering organisation, we are looking for leaders who can build strong teams, uphold high standards, and deliver reliably at pace.
Role Responsibilities
- Deploy and Manage Kubernetes Clusters, deployed at scale to support AI centric workloads, across both our bare metal clusters and via trusted partner infrastructure
- Develop Kubernetes Manifests and Operators: Facilitate application deployments and maintain Kubernetes-native services for networking, storage, security, identity and infrastructure management
- Optimize Linux system configuration including kernel, driver, filesystem and services to support workloads running via our orchestration layer
- Build and maintain automation scripts and infrastructure as code to support platform lifecycle, as well as simplifying troubleshooting for Incident resolution and provision of tooling for our support organisation
- Apply ITSM frameworks: Incident, Major Incident, Change Management, and service improvement.
- Maintain and enhance Radiant’s observability stack: Prometheus, Grafana, and custom monitoring integrations
- Operate and support services in 24x7 production environments, including on-call rotation
- Contribute to Incident postmortem analyses, root cause analysis, document learnings, and automate remediations
- Mentor junior engineers and act as an Operational requirements consultant to other departments
- Communicate technical decisions clearly to non-technical stakeholders and customers
- Uphold a culture of: do, document, automate
- Willingness to cross train with Platform Engineering/Platform SRE to fully support both our infrastructure and platform stacks.
- Willingness to cross train with HPC Engineering, supported by NVIDIA to enhance our HPC supportability offering
Requirements
- 5+ Years Proven experience in globally scaled, performance-intensive environments operating to a 24/7 support model in an SRE or equivalent role
- 3+ years experience in both running, deploying and optimising orchestration platforms with a strong emphasis on Kubernetes
- Expert-level Linux administration, especially Ubuntu distributions
- Proficiency in system tuning, disk I/O optimization, and hardware-level performance tweaks
- Strong networking fundamentals: TCP/IP, DNS, DHCP, VLANs, routing, switching
- Strong experience with API interrogation
- Strong experience with infrastructure scripting and automation (Bash, Python, Ansible)
- Deep understanding of observability principles and tools (Prometheus, Grafana preferred)
- Strong grasp of ITSM and service operation best practices
- Excellent communication and mentorship skills
- Comfortable interfacing with internal stakeholders and external customers
- Bonus: Knowledge of running AI workloads via orchestration platforms
Bonus Requirements
- Bachelor or Masters Level degree in Computer Science, Engineering or related field, or equivalent experience.
- LPIC Certifications
- ITIL Foundation level qualification or equivalent experience
- Certified Kubernetes Administrator (CKA)
Qualities we look for:
- You approach problems with a systems mindset - balancing practical execution with long-term scalability
- You elevate the team, setting high standards for technical quality and engineering excellence.
- You hold yourself and others accountable - giving direct feedback and expecting the same
- You take initiative, owning challenges end-to-end and proactively driving solutions.
2,075,533 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →