HPC Engineer
AlphaGrep Securities
| Company | AlphaGrep Securities |
| Category | Engineering |
| Location | Shanghai |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 15 Jun 2023 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (greenhouse) |
Description
关于我 们 / About Us
AlphaGrep 是一家全球领先的量化交易公司,专注于股票、商品、外汇及固定收益等资产的算法交易。我们在国际市场拥有显著份额,依托自主开发的超低延迟系统与严格的风控体系,持续构建高效能策略。
AlphaGrep is a leading global quantitative trading firm specializing in algorithmic strategies across equities, commodities, FX, and fixed income. We hold significant market share internationally, powered by proprietary low-latency infrastructure and robust risk controls.
AlphaGrep China 是专注于中国市场人民币资产管理的机构,服务对象涵盖机构投资人、家族办公室与高净值客户。我们深耕中国资本市场,涵盖股票及衍生品等多类资产,结合全球量化研究体系与本地实战经验,构建多元交易策略,致力于实现长期稳健增长。
AlphaGrep China is a dedicated RMB asset management platform focused on the Chinese market, serving institutional investors, family offices, and high-net-worth individuals. We leverage deep expertise in China’s equity and derivatives markets, together with AlphaGrep’s global quantitative research capabilities, to deliver diversified strategies targeting long-term and stable returns.
岗位职责 / Responsibilities
我们正在寻找一位 HPC / GPU 集群工程师,协助设计、运营并持续优化公司用于分布式模型训练及高性能计算的大规模 GPU 计算环境。您将端到端负责集群的性能与稳定性——从 GPU 硬件与互联网络,到存储层、调度系统、监控平台,以及保障数百块加速器高效运行的全套工具链。
We are looking for an HPC / GPU Cluster Engineer to help design, operate, and continuously optimize our large-scale GPU compute environment used for distributed model training and high-performance workloads. You will own the performance and reliability of the cluster end to end — from the GPUs and interconnect fabric up through the storage layer, scheduler, monitoring, and the tooling that keeps hundreds of accelerators running efficiently.
任职要求 / Qualifications
具备生产环境下大规模 GPU 或 HPC 集群的运营经验,能够系统性地识别并解决性能瓶颈。
Experience operating large GPU or HPC clusters in production, identifying performance bottlenecks and resolving them systematically.
具备 RDMA 编程及底层调试的实战经验,包括 RDMA verbs / libibverbs、内存注册、队列对、完成队列、RDMA 读写、发送/接收操作及性能基准测试。
Hands-on experience with RDMA programming and low-level RDMA debugging, including RDMA verbs / libibverbs, memory registration, queue pairs, completion queues, RDMA read/write, send/recv, and performance benchmarking.
熟悉分布式训练的核心底层技术,包括 InfiniBand 和/或 RoCEv2、GPUDirect RDMA、GPUDirect Storage、NCCL、CUDA 驱动及 OFED。
Practical experience with the technologies underpinning distributed training, including InfiniBand and/or RoCEv2, GPUDirect RDMA, GPUDirect Storage, NCCL, CUDA drivers, and OFED.
具备 Slurm 等调度系统的使用经验,包括安装配置、队列/分区管理、作业故障排查及监控。
Experience with workload schedulers such as Slurm, including setup, configuration, queue/partition management, job troubleshooting, and monitoring.
具备高性能共享存储或并行/分布式文件系统的搭建与管理经验,如 Lustre、BeeGFS、WEKA、VAST、DDN/ExaScaler 等。
Experience setting up and managing high-performance shared storage or parallel/distributed filesystems such as Lustre, BeeGFS, WEKA, VAST, DDN/ExaScaler, or similar systems.
熟练掌握 Python、Bash,优先具备 C/C++ 能力,用于集群自动化、诊断、基准测试及监控开发。
Solid scripting/programming ability in Python, Bash, and preferably C/C++, f
1,014,484 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →