Staff / Senior Software DevOps Engineer
Olix
| Company | Olix |
| Category | Uncategorised |
| Location | Toronto |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 15 Jul 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (ashby) |
Description
ABOUT OLIX
AI is growing faster than any technology in history and the explosion in demand has created a massive infrastructure gap; we can no longer build chips or power stations fast enough to keep up. The industry is still leaning on a ten-year-old hardware blueprint that has reached its limit. A new paradigm that is faster and more efficient will be the biggest economic opportunity of the next century and create the most important company of the next decade. The OLIX Decode Accelerator 1 (DX-1) is the first accelerator architected specifically for decode. Rack-scale co-design of logic, data movement, packaging, optics and interconnect enables a step change in system level performance.
THE ROLE
We're searching for a Staff/Senior Software DevOps Engineer to own the build, test, and CI flows that the entire DX-1 software stack, including the compiler, runtime, simulator, and framework integration, depends on. Ours is a large test suite that asserts token-exact correctness against golden references, and it has to run across scarce, expensive resources that span both simulation compute and hardware-in-the-loop testing, including simulator and emulator boxes alongside DX-1 and prototype-platform boards. Your mission is to keep that system fast, trustworthy, observable, and affordable as the test suite, the team, and the resource pool all grow.
This is a build-and-test role, not product-serving SRE. You'll work where CI, the runner fleet, and the test hardware meet, partnering closely with the infrastructure, compiler, runtime, simulator, and modelling teams. At the Senior/Staff level, your impact is the velocity of every engineer who depends on this system: how fast they get a trustworthy signal, how rarely they wait on a machine or a flaky run, and how much they can self-serve without coming to you. That leverage, through the standards, platforms, and shared resource model others build on, is what we're hiring for far more than any single system you ship.
RESPONSIBILITIES
Own the Build & Test Pipelines: Design, build, and own CI pipelines and test execution across PR, merge, and nightly lanes that gate the entire software stack, balancing fast feedback with coverage and cost.
Scale Test Execution: Split a large, slow suite into staged lanes, parallelize it with real test isolation, and cache aggressively using content-addressed keys so feedback stays fast and cost-effective as the suite and the team grow, rather than relying on simply adding more machines.
Manage the Fleet & Scarce Resources: Run CI across a heterogeneous fleet of cloud and self-hosted machines, and give the team fair, monitored, fail-fast shared access to scarce and expensive hardware, keeping it reliable, well utilized, and never a silent bottleneck.
Build the Performance & Readiness Signal: Stand up performance regression baselines the team trusts using pinned hardware, rolling baselines, sound metric aggregation, and deterministic testing. Turn CI and test signals into CI health and product readiness dashboards that drive real decisions.
Own Software Observability: Choose the metrics store that scales to many time series across daily runs with long-lived history, making dashboards for observable software.
Set Standards: Define the flows that keep builds and test runs hermetic and reproducible, and make the system fail closed while containing the blast radius when something is misconfigured or a job is untrusted.
SKILLS & EXPERIENCE
- Experience in build/test infrastructure, CI/CD, developer productivity, or large-scale systems and release engineering, with demonstrated end-to-end ownership of a large test or CI system
- Experience scaling a large test suite through staged lanes, parallelism with real isolation, content-addressed caching, and maintaining fast, cost-effective feedback as the suite grows
- Experience managing heterogeneous CI runner fleets across cloud and on-prem environments, including VMs, containers, and bare-met
991,236 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →