Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

HPC Engineer

AlphaGrep Securities
CompanyAlphaGrep Securities
CategoryEngineering
LocationShanghai
RemoteOn-site (inferred)
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted9 Jun 2026
Last verified12 Aug 2026
SourceThe employer's own careers page (company_site)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
AlphaGrep is a leading global quantitative trading firm specializing in algorithmic strategies across equities, commodities, FX, and fixed income. This role is for an HPC/GPU Cluster Engineer who will design, operate, and continuously optimize large-scale GPU compute environments used for distributed model training and high-performance computing workloads. What You'll Do • Own end-to-end performance and reliability of GPU clusters, spanning GPU hardware, interconnect fabric, storage layer, scheduler, and monitoring systems • Identify and resolve performance bottlenecks systematically in production HPC and GPU cluster environments • Configure, manage, and troubleshoot workload schedulers such as Slurm, including queue/partition management and job diagnostics • Design and manage high-performance shared storage and parallel/distributed file systems • Develop cluster automation, diagnostic tools, benchmarking suites, and monitoring solutions using Python, Bash, and C/C++ What You Need • Production experience operating large-scale GPU or HPC clusters with proven ability to identify and resolve performance bottlenecks • Hands-on RDMA programming and debugging experience including RDMA verbs/libibverbs, memory registration, queue pairs, completion queues, and performance benchmarking • Practical knowledge of distributed training technologies including InfiniBand and/or RoCEv2, GPUDirect RDMA, GPUDirect Storage, NCCL, CUDA drivers, and OFED • Strong Linux system administration skills including networking, filesystems, kernel/driver issues, process/resource management, and performance debugging • Proficiency in Python and Bash scripting; C/C++ skills preferred Nice to Have • C/C++ programming capability for advanced cluster automation and tooling • Deep familiarity with HPC networking design including blocking/non-blocking fabrics, topology, congestion control, and end-to-end troubleshooting Health and wellness programs including gym membership, stocked kitchen, and generous vacation