HPC Engineer
AlphaGrep Securities
| Company | AlphaGrep Securities |
| Category | Engineering |
| Location | Shanghai |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 9 Jun 2026 |
| Last verified | 12 Aug 2026 |
| Source | The employer's own careers page (company_site) |
Description
AlphaGrep is a leading global quantitative trading firm specializing in algorithmic strategies across equities, commodities, FX, and fixed income. This role is for an HPC/GPU Cluster Engineer who will design, operate, and continuously optimize large-scale GPU compute environments used for distributed model training and high-performance computing workloads.
What You'll Do
• Own end-to-end performance and reliability of GPU clusters, spanning GPU hardware, interconnect fabric, storage layer, scheduler, and monitoring systems
• Identify and resolve performance bottlenecks systematically in production HPC and GPU cluster environments
• Configure, manage, and troubleshoot workload schedulers such as Slurm, including queue/partition management and job diagnostics
• Design and manage high-performance shared storage and parallel/distributed file systems
• Develop cluster automation, diagnostic tools, benchmarking suites, and monitoring solutions using Python, Bash, and C/C++
What You Need
• Production experience operating large-scale GPU or HPC clusters with proven ability to identify and resolve performance bottlenecks
• Hands-on RDMA programming and debugging experience including RDMA verbs/libibverbs, memory registration, queue pairs, completion queues, and performance benchmarking
• Practical knowledge of distributed training technologies including InfiniBand and/or RoCEv2, GPUDirect RDMA, GPUDirect Storage, NCCL, CUDA drivers, and OFED
• Strong Linux system administration skills including networking, filesystems, kernel/driver issues, process/resource management, and performance debugging
• Proficiency in Python and Bash scripting; C/C++ skills preferred
Nice to Have
• C/C++ programming capability for advanced cluster automation and tooling
• Deep familiarity with HPC networking design including blocking/non-blocking fabrics, topology, congestion control, and end-to-end troubleshooting
Health and wellness programs including gym membership, stocked kitchen, and generous vacation