DevOps Engineer, Cloud Infrastructure & Live Games
Big Viking Games
| Company | Big Viking Games |
| Category | Engineering |
| Location | Toronto |
| Remote | On-site (inferred) |
| Employment | Full-time |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 15 Jul 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (workable) |
Description
About Big Viking Games Big Viking Games is a Canadian gaming company focused on building, operating, and growing long-standing online game communities. Our games have entertained players for years, supported by loyal audiences, live operations, evolving content systems, product innovation, and deep player-driven economies. Our flagship titles, YoWorld and FishWorld, have served millions of players over their lifetime. These are enduring live-service virtual worlds with rich in-game economies, virtual goods, social interaction, and long-term player engagement at their core. We are entering a new phase of modernization and growth, with a focus on stronger infrastructure, better automation, practical AI adoption, improved reliability, stronger security practices, and scalable systems that help our games and teams perform at a higher level. About the Role Big Viking Games is hiring a Senior DevOps Engineer to help design, maintain, secure, and modernize the infrastructure that supports our live-service games and internal development workflows. This is a hands-on role for someone who understands cloud infrastructure, automation, CI/CD, containers, monitoring, uptime, and production reliability — and who is comfortable working with legacy production systems alongside modern infrastructure patterns. Our games have been running for over a decade; the infrastructure reflects that history, and the right person sees that as an interesting challenge rather than a dealbreaker. You will work closely with engineering, product, QA, data, and live operations teams to improve how we build, deploy, monitor, and operate our systems. The right person is practical, security-aware, automation-minded, and able to balance speed with reliability. They can own infrastructure in a live production environment, improve DevOps processes, reduce manual work, and help development teams ship safely and efficiently. This is a hybrid role based in Toronto, with an expectation of working in office three days per week. Live-service games require operational awareness outside regular business hours, including periodic on-call and incident response availability. Requirements What You'll Do Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and related platforms that support our live games, data systems, internal tools, and AI-powered operational workflows. Drive infrastructure modernization while maintaining uptime for live games with active player communities — every improvement ships while the plane is flying. Build, maintain, and improve automation for deployments, environment management, provisioning, secrets rotation, and operational workflows — reducing manual toil and human error. Implement and maintain Infrastructure as Code using tools such as Terraform, CloudFormation, CDK, or similar technologies. Maintain and monitor data pipelines between game source databases (MariaDB), the Snowflake data warehouse, and downstream analytics and reporting systems — ensuring pipeline health, freshness, and alerting when data stops flowing. Improve CI/CD pipelines, release workflows, and deployment reliability so development teams can ship safely and frequently. Own secrets and credential lifecycle management across platforms — including API key rotation, access controls, environment variable governance, and least-privilege practices. Support and improve the infrastructure that powers AI and automation tooling, including API integrations, MCP servers, serverless functions, webhook reliability, and orchestration platforms. Improve observability across the stack: logging, metrics, alerting, dashboards, and operational visibility — with particular attention to early detection of silent failures in data pipelines and production systems. Support incident response, root cause analysis, remediation planning, and post-incident improvements. Help manage cloud spend, infrastructure usage, resource tagging, and e
991,236 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →