Spécialiste CHP
DRW Montreal
| Company | DRW Montreal |
| Category | Operations & Admin |
| Location | Montreal |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Mid |
| Salary | Not stated by the employer |
| Posted | 5 May 2026 |
| Last verified | 12 Aug 2026 |
| Source | The employer's own careers page (company_site) |
Description
DRW is a diversified trading firm with over 30 years of experience operating across global markets. This role is a GPU Infrastructure Specialist joining the AI and multi-asset systematic strategies team, responsible for designing and operating GPU infrastructure that powers AI and machine learning workloads across the entire infrastructure stack, from bare-metal hardware to model deployment.
What You'll Do
• Deploy, maintain, and optimize GPU infrastructure for large-scale LLM inference workloads, including provisioning, configuration, and deployment of GPU server fleets
• Design and implement distributed model serving solutions for multi-node and multi-GPU deployments
• Manage Kubernetes clusters with GPU support for LLM and ML workloads
• Configure network infrastructure including load balancers, firewalls, and inter-node communication for GPU clusters
• Diagnose performance bottlenecks at all levels: hardware, drivers, network, and application layer
• Collaborate with ML engineers to profile model performance and implement inference acceleration techniques
• Improve reliability through monitoring, alerting, capacity planning, and incident management
What You Need
• Bachelor's or Master's degree in Computer Science, Systems Engineering, or related field
• 5+ years of experience in DevOps, SRE, or infrastructure engineering
• Strong experience with GPU infrastructure, model serving frameworks (vLLM, SGLang), and GPU driver management
• Hands-on experience optimizing deep learning workloads (inference or training) on GPU clusters
• In-depth knowledge of Linux systems, including network configuration, storage optimization, and Kubernetes orchestration
• Proficiency in Python and Bash scripting for automation
• Solid understanding of distributed systems, network protocols (TCP/IP, HTTP/2), and load balancing
Nice to Have
• Experience with infrastructure-as-code tools (Ansible, Terraform, or equivalent)
• Experience with monitoring and observability tools (Prometheus, Grafana, or equivalent)
Recognized as one of Canada's best employers for 8 consecutive years. Comprehensive benefits package, commitment to continuous training and development, employee wellness and work-life balance focus, community initiatives and volunteering programs. See https://drw.com/fr/work-at-drw/avantages-montreal for full details.