Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Junior HPC Systems Engineer

Parallel Works
CompanyParallel Works
CategoryUncategorised
LocationChicago
RemoteRemote
EmploymentFull-time
LevelNot stated
SalaryNot stated by the employer
Posted1 Aug 2026
Last verified3 Aug 2026
SourceEmployer career page (workable)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
About Parallel Works Parallel Works builds and operates ACTIVATE, a control plane for high performance computing and AI. Our customers run large scientific and AI workloads across their own on-premises clusters, Government and commercial cloud, and commercial GPU providers, and ACTIVATE gives them one way in to all of it. The high security boundary is authorized at Impact Level 5, with FIPS validated cryptography and STIG hardening throughout. The work reaches most fields that depend on computing at scale: weather and climate forecasting, defense and intelligence programs, aerospace and structural analysis, molecular and materials science, energy, and AI research. A quarter here can include standing up a GPU cluster for one of those communities, federating a laboratory's existing on-premises system with burst capacity it did not have before, and getting a domain code written decades ago to run on current hardware. Customer success sets our priorities. We are a small engineering company, so engineers here work directly with the people using the systems and carry a problem from the first report through to the fix. This is what we call mission engineering: understanding what a customer is trying to accomplish and why the computing matters to it. About the role Parallel Works is hiring a Junior HPC Systems Engineer to learn supercomputing operations on production systems. The role starts with monitoring, node health, and account and allocation work, and moves into cluster builds and escalations. The systems are hybrid: on-premises clusters the customer owns, accredited Government cloud regions, and commercial GPU providers, often in the same day. We are looking for strong Linux fundamentals and interest in how large systems behave. Running a supercomputer is not a prerequisite. The role leads into senior systems engineering, with cluster builds handled independently after about a year. What you will do Monitor and respond:  cluster health, node state, queue behavior, and alerting, with first action on node failures, stuck jobs, and filesystem alerts. Accounts and allocations:  users, groups, Slurm accounts, and allocations, kept consistent across venues. Node lifecycle:  health checks on GPU and CPU nodes, draining and returning nodes, and the escalation path on faults, which is site staff or a vendor on customer hardware and a support case on cloud capacity. Automation:  extend Ansible playbooks and operational scripts, and replace the manual steps you find yourself repeating. Patching and hardening:  patches and hardening baselines under senior review, plus scan evidence for the security package. Documentation and on call:  keep runbooks current, and join the on call rotation once trained into it.
HOUSE ADYou found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →