Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

DevOps Engineer, Agentic Operation & Live Games

Big Viking Games
CompanyBig Viking Games
CategoryEngineering
LocationToronto
RemoteOn-site (inferred)
EmploymentFull-time
LevelNot stated
SalaryNot stated by the employer
Posted15 Jul 2026
Last verified11 Aug 2026
SourceEmployer ATS (workable)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
About Big Viking Games Big Viking Games is a Canadian gaming company focused on building, operating, and growing long-standing online game communities. Our games have entertained players for years, supported by loyal audiences, live operations, evolving content systems, product innovation, and deep player-driven economies. Our flagship titles, YoWorld and FishWorld, have served millions of players over their lifetime. These are enduring live-service virtual worlds with rich in-game economies, virtual goods, social interaction, and long-term player engagement at their core. We are entering a new phase of modernization and growth, with a focus on stronger infrastructure, better automation, practical AI adoption, improved reliability, stronger security practices, and scalable systems that help our games and teams perform at a higher level. About the Role Big Viking Games is hiring a Senior DevOps Engineer to modernize the infrastructure behind our live-service games — and to do it by building automation that thinks, not just scripts that run. Our games run on mature tech stacks with large data volumes and player counts — real systems, with real players on them, where the constraints are genuine and the consequences are visible. There is meaningful room to automate how they are operated, and we want someone who builds that: tooling that takes routine work off people's hands and runs safely against production without compromising uptime or data integrity. This is a hands-on senior role for an engineer who is fluent in modern cloud infrastructure — AWS, containers, IaC, CI/CD, observability, production reliability — and who has started using agentic coding tools and tool-calling systems to do infrastructure work that previously required a person. If you have wired a coding agent into your CI/CD, built an MCP server so tooling could act on your infrastructure safely, or replaced a manual runbook with something that diagnoses and remediates on its own, this role is aimed at you. The defining trait is self-direction , and it works in two directions. Given a backlog, you improve on how the work gets done — bringing agentic approaches to problems that were scoped as manual ones, and solving the underlying issue rather than the individual ticket. Left to your own judgement, you find the repetitive work nobody has flagged, decide what's worth automating, and build it. In both cases you own the guardrails that make automation safe to run against production. This is a hybrid role based in Toronto, with an expectation of working in office three days per week. Live-service games require operational awareness outside regular business hours, including periodic on-call and incident response availability. Requirements What You'll Do Build agentic automation for infrastructure work Design, build, and operate agentic tooling that performs real DevOps work — diagnosis, remediation, provisioning, routine maintenance — rather than only summarizing or suggesting. Build and maintain the integration layer that lets tooling act safely on our systems: MCP servers, API integrations, webhook-driven orchestration, serverless functions, and the permissions and audit trails around them. Convert manual runbooks, SOPs, and recurring operational chores into automation that runs unattended, with sensible escalation when it shouldn't proceed alone. Establish the guardrails that make automated action against production defensible: least-privilege scopes, dry-run and approval paths for destructive operations, logging of what automation did and why, and clear rollback. Use agentic coding tools to accelerate your own infrastructure work — IaC authoring, migration scripting, log and incident analysis, documentation — and improve how the wider engineering team does the same. Own and modernize live production infrastructure Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and