Senior HPC Systems Engineer
Parallel Works
| Company | Parallel Works |
| Category | Uncategorised |
| Location | United States |
| Remote | Remote |
| Employment | Full-time |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 1 Aug 2026 |
| Last verified | 3 Aug 2026 |
| Source | Employer career page (workable) |
Description
About Parallel Works Parallel Works builds and operates ACTIVATE, a control plane for high performance computing and AI. Our customers run large scientific and AI workloads across their own on-premises clusters, Government and commercial cloud, and commercial GPU providers, and ACTIVATE gives them one way in to all of it. The high security boundary is authorized at Impact Level 5, with FIPS validated cryptography and STIG hardening throughout. The work reaches most fields that depend on computing at scale: weather and climate forecasting, defense and intelligence programs, aerospace and structural analysis, molecular and materials science, energy, and AI research. A quarter here can include standing up a GPU cluster for one of those communities, federating a laboratory's existing on-premises system with burst capacity it did not have before, and getting a domain code written decades ago to run on current hardware. Customer success sets our priorities. We are a small engineering company, so engineers here work directly with the people using the systems and carry a problem from the first report through to the fix. This is what we call mission engineering: understanding what a customer is trying to accomplish and why the computing matters to it. About the role Parallel Works is hiring a Senior HPC Systems Engineer to build and run the clusters behind our defense and research programs. The work covers GPU node bring-up, Slurm configuration, fabric and storage troubleshooting, security hardening, and Tier 3 escalation. The computing environments are hybrid. Some clusters are customer owned hardware on site, some run in accredited Government cloud regions, and some are dedicated GPU clusters at commercial providers. On several programs the on-premises systems carry the primary load and cloud takes the overflow. The position is senior: it handles the escalations the rest of the team cannot resolve, and it trains the junior engineers. What you will do Cluster operations: build and operate production Slurm clusters. slurmctld and slurmdbd, partitions and QOS, accounts and fair share, GPU GRES, prolog and epilog, cgroup enforcement. Hybrid federation: connect customer owned clusters to the control plane, reconciling their site scheduler, storage, and identity source so accounts and allocations behave the same in every venue. On-premises hardware: bare metal provisioning, out of band management, firmware, rack networking, and fault coordination with site staff or vendors. GPU and fabric: validate GPU nodes before users arrive. Driver and CUDA stack, DCGM health checks, XID triage, fabric manager and NVLink checks, InfiniBand verification, NCCL tuning. Storage and automation: tune parallel and high throughput storage, and write the Ansible, Terraform, and image build pipelines that make a cluster reproducible. Security and escalation: STIG hardening, scan remediation, FIPS validated cryptography, security package artifacts, Tier 3 escalations, and a share of the on call rotation.
1,970,675 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →