Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Senior Systems Engineer, Test Frameworks & Validation Platform

CoreWeave
CompanyCoreWeave
CategoryUncategorised
LocationLivingston
RemoteOn-site (inferred)
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted29 Jul 2026
Last verified30 Jul 2026
SourceEmployer career page (greenhouse)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at  www.coreweave.com .   What You'll Do: The Systems Engineering team owns the host software stack that turns a freshly provisioned bare-metal machine into a healthy Kubernetes worker — the OS image, kernel and drivers, and the hardware and firmware beneath them. We build and operate the test framework that qualifies every server image on real physical hardware before it reaches one of the largest GPU fleets in the world, and we're extending that framework into HPC verification, Slurm-on-Kubernetes, and further down the stack. About the role: As a Senior Software Engineer on the Systems Engineering team, you will own and evolve our test framework — in-house, Kubernetes-native system that qualifies host images and infrastructure on real hardware. You'll harden the framework's core, broaden what it can validate, and keep CI fast and trustworthy at fleet scale, where a 0.1% failure rate is thousands of GPUs. Day-to-day, you'll build test harnesses and CI infrastructure, extend coverage into HPC and Slurm-on-Kubernetes (SUNK), and shape the abstractions other engineers use to add their own tests without reinventing the framework underneath them. Some of what you'll work on: Own and extend the test framework — that pins to a real host, boots the configuration under test, runs a containerized check, validates it, and records a structured result. Broaden what the framework qualifies: HPC and fabric verification, Slurm-on-Kubernetes (SUNK), and further down the stack into firmware and hardware, building the abstractions that make new coverage easy to add. Keep the pipeline fast, hermetic, and trustworthy — boot consistency, per-host resource locking, and ruthless flaky-test elimination — so engineers trust the signal. Own the results and reporting path end to end: structured test results, storage, dashboards, and the triage surfaces engineers use to turn a red run into a root cause quickly. Explore AI-native testing — LLM-driven log triage, failure classification, and regression detection — as the fleet and test surface grow. Collaborate with firmware, kernel, imaging, and HPC teams to embed testing into their release process. Who You Are: 3+ years of experience building test infrastructure, systems software, or platform tooling at scale, with real ownership of frameworks or automation for low-level software. Fluent in Python, with either proven Rust experience or a strong systems background (Go, C, C++) and a genuine appetite to build in Rust — the framework's core is Rust, and you'll live in it. Comfortable operating in a Kubernetes environment and reasoning about how software is built, containerized, deployed, and tested. Solid Linux systems background — the boot chain, kernel and drivers, low-level debugging — and comfort working close to the hardware. Real testing discipline: you've built automation that proves systems work, and you have strong opinions about flakiness, hermeticity, and signal. Clear communicator who treats the test framework as a product other engineers want to use, not a chore they route around. Preferred: Rust and Kubernetes-native workflow orchestration (e.g., Argo Workflows) experience. HPC or large-cluster experience — InfiniBand/RoCE, GPU/accelerator validation, or performance-regression frameworks. Slurm or Slurm-on-Kubernetes (SUNK) experience. Firmware or lower-level hardware validation experie
HOUSE ADYou found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →