Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

HPC Engineer

AlphaGrep Securities
CompanyAlphaGrep Securities
CategoryEngineering
LocationShanghai
RemoteOn-site (inferred)
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted15 Jun 2023
Last verified30 Jul 2026
SourceEmployer career page (greenhouse)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
关于我 们 / About Us AlphaGrep 是一家全球领先的量化交易公司,专注于股票、商品、外汇及固定收益等资产的算法交易。我们在国际市场拥有显著份额,依托自主开发的超低延迟系统与严格的风控体系,持续构建高效能策略。 AlphaGrep is a leading global quantitative trading firm specializing in algorithmic strategies across equities, commodities, FX, and fixed income. We hold significant market share internationally, powered by proprietary low-latency infrastructure and robust risk controls. AlphaGrep China 是专注于中国市场人民币资产管理的机构,服务对象涵盖机构投资人、家族办公室与高净值客户。我们深耕中国资本市场,涵盖股票及衍生品等多类资产,结合全球量化研究体系与本地实战经验,构建多元交易策略,致力于实现长期稳健增长。 AlphaGrep China is a dedicated RMB asset management platform focused on the Chinese market, serving institutional investors, family offices, and high-net-worth individuals. We leverage deep expertise in China’s equity and derivatives markets, together with AlphaGrep’s global quantitative research capabilities, to deliver diversified strategies targeting long-term and stable returns.   岗位职责 / Responsibilities 我们正在寻找一位 HPC / GPU 集群工程师,协助设计、运营并持续优化公司用于分布式模型训练及高性能计算的大规模 GPU 计算环境。您将端到端负责集群的性能与稳定性——从 GPU 硬件与互联网络,到存储层、调度系统、监控平台,以及保障数百块加速器高效运行的全套工具链。 We are looking for an HPC / GPU Cluster Engineer to help design, operate, and continuously optimize our large-scale GPU compute environment used for distributed model training and high-performance workloads. You will own the performance and reliability of the cluster end to end — from the GPUs and interconnect fabric up through the storage layer, scheduler, monitoring, and the tooling that keeps hundreds of accelerators running efficiently.   任职要求 / Qualifications 具备生产环境下大规模 GPU 或 HPC 集群的运营经验,能够系统性地识别并解决性能瓶颈。 Experience operating large GPU or HPC clusters in production, identifying performance bottlenecks and resolving them systematically. 具备 RDMA 编程及底层调试的实战经验,包括 RDMA verbs / libibverbs、内存注册、队列对、完成队列、RDMA 读写、发送/接收操作及性能基准测试。 Hands-on experience with RDMA programming and low-level RDMA debugging, including RDMA verbs / libibverbs, memory registration, queue pairs, completion queues, RDMA read/write, send/recv, and performance benchmarking. 熟悉分布式训练的核心底层技术,包括 InfiniBand 和/或 RoCEv2、GPUDirect RDMA、GPUDirect Storage、NCCL、CUDA 驱动及 OFED。 Practical experience with the technologies underpinning distributed training, including InfiniBand and/or RoCEv2, GPUDirect RDMA, GPUDirect Storage, NCCL, CUDA drivers, and OFED. 具备 Slurm 等调度系统的使用经验,包括安装配置、队列/分区管理、作业故障排查及监控。 Experience with workload schedulers such as Slurm, including setup, configuration, queue/partition management, job troubleshooting, and monitoring. 具备高性能共享存储或并行/分布式文件系统的搭建与管理经验,如 Lustre、BeeGFS、WEKA、VAST、DDN/ExaScaler 等。 Experience setting up and managing high-performance shared storage or parallel/distributed filesystems such as Lustre, BeeGFS, WEKA, VAST, DDN/ExaScaler, or similar systems. 熟练掌握 Python、Bash,优先具备 C/C++ 能力,用于集群自动化、诊断、基准测试及监控开发。 Solid scripting/programming ability in Python, Bash, and preferably C/C++, f
HOUSE AD1,014,484 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →