DevOps Engineer
Skillonnet 1744715103
| Company | Skillonnet 1744715103 |
| Category | Engineering |
| Location | — |
| Remote | — |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 11 Jun 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (teamtailor) |
Description
Company Description We are SkillOnNet, leading the igaming entertainment by providing our customers with the most entertaining and trustworthy experience possible, while also reinventing the gambling industry. We are home to more than 30 well-known brands, including PlayOJO, DruckGluck, BacanaPlay, Genting, and many more. We are committed to long-term development and sustainability, and we are trying to revolutionize our industry for the benefit of our players, ourselves, and the entertainment industry as a whole. Job Description We are seeking a skilled and proactive DevOps Engineer to join our growing Tech team. This role is critical in ensuring the reliability, scalability, and performance of our infrastructure and development pipelines. You will work closely with the Heads of Development, System Administrators, and the Information Security team to maintain and evolve our systems and services. Responsibilities include: Kubernetes (K8s) Administration: Deploy, manage, scale, and troubleshoot Kubernetes clusters, ensuring high availability, performance, and reliability across production and non-production environments. Observability & Monitoring: Design, maintain, and upgrade monitoring and alerting platforms using Grafana, Prometheus, and Alert manager to improve system visibility, incident detection, and operational efficiency. Containerization & Image Management: Build, optimize, secure, and maintain Docker images and containerized applications, following best practices for performance, scalability, and vulnerability management. CI/CD Automation: Design and maintain GitHub Actions workflows to automate testing, integration, deployment, and release processes, reducing manual intervention and accelerating software delivery. Infrastructure Operations: Manage Linux servers and infrastructure components, including performance tuning, capacity planning, backup strategies, disaster recovery, and system hardening. Ceph Storage Administration: Deploy, maintain, and optimize Ceph distributed storage clusters, ensuring data durability, scalability, and high availability for critical workloads. Cloudflare Platform Management: Configure and manage Cloudflare services, including DNS, WAF, CDN, SSL/TLS, caching policies, and Workers to enhance security, performance, and application availability. Security & Reliability Engineering: Implement infrastructure security best practices, automate compliance controls, manage secrets, and improve system resilience through proactive monitoring and incident response. Infrastructure as Code (IaC): Automate infrastructure provisioning and configuration management using tools such as Terraform and Ansible to ensure consistency, repeatability, and scalability. Incident Management & Troubleshooting: Lead root-cause analysis, resolve production issues, and implement preventive measures to minimize downtime and improve platform stability. What we are looking for: Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience. Proven experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Operations roles. Strong hands-on expertise with Kubernetes and Docker, including deploying, scaling, troubleshooting, and maintaining containerized workloads in production environments. Experience designing and managing CI/CD pipelines, preferably using GitHub Actions, with a focus on automation, reliability, and deployment best practices. Solid knowledge of observability and monitoring platforms, including Grafana, Prometheus, and Alertmanager, with experience implementing effective monitoring and alerting strategies. Experience administering distributed storage systems such as Ceph, including performance optimization, capacity planning, and high-availability configurations. Strong understanding of