Member of Technical Staff — Compute Cluster
Causal
| Company | Causal |
| Category | Engineering |
| Location | San Francisco |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Mid |
| Salary | Not stated by the employer |
| Posted | 19 Jul 2026 |
| Last verified | 12 Aug 2026 |
| Source | The employer's own careers page (company_site) |
Description
This AI research company is building large physics foundation models for causal intelligence and weather prediction. You will design, build, and operate large-scale GPU clusters that power training, evaluation, and serving infrastructure for the research team.
What You'll Do
• Design, deploy, and operate large distributed GPU clusters end-to-end, including provisioning, imaging, upgrades, and capacity planning
• Extend scheduling and orchestration systems such as Kubernetes and Slurm for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads
• Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers
• Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage
• Monitor and continuously improve reliability and error recovery; build observability to catch failures before researchers do
• Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs
What You Need
• Experience operating large-scale GPU clusters and container orchestration frameworks such as Kubernetes, Slurm, and Docker
• Strong systems background in Linux, networking, storage, and infrastructure-as-code
• Knowledge of cloud platforms including GCP, AWS, or Azure and their ML/AI service offerings
• Understanding of monitoring, logging, observability, and version control best practices for ML systems
• Familiarity with CUDA and NCCL, and performance profiling for distributed workloads
• Ability to own deliverables end-to-end from requirements through autonomous execution