Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Member of Technical Staff — Compute Cluster

Causal
CompanyCausal
CategoryEngineering
LocationSan Francisco
RemoteOn-site (inferred)
EmploymentNot stated
LevelMid
SalaryNot stated by the employer
Posted19 Jul 2026
Last verified12 Aug 2026
SourceThe employer's own careers page (company_site)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
This AI research company is building large physics foundation models for causal intelligence and weather prediction. You will design, build, and operate large-scale GPU clusters that power training, evaluation, and serving infrastructure for the research team. What You'll Do • Design, deploy, and operate large distributed GPU clusters end-to-end, including provisioning, imaging, upgrades, and capacity planning • Extend scheduling and orchestration systems such as Kubernetes and Slurm for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads • Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers • Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage • Monitor and continuously improve reliability and error recovery; build observability to catch failures before researchers do • Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs What You Need • Experience operating large-scale GPU clusters and container orchestration frameworks such as Kubernetes, Slurm, and Docker • Strong systems background in Linux, networking, storage, and infrastructure-as-code • Knowledge of cloud platforms including GCP, AWS, or Azure and their ML/AI service offerings • Understanding of monitoring, logging, observability, and version control best practices for ML systems • Familiarity with CUDA and NCCL, and performance profiling for distributed workloads • Ability to own deliverables end-to-end from requirements through autonomous execution