HPC Engineer, Metal Net
CoreWeave Europe
| Company | CoreWeave Europe |
| Category | Engineering |
| Location | Warsaw |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 20 May 2026 |
| Last verified | 3 Aug 2026 |
| Source | Employer ATS (greenhouse) |
Description
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com .
We're proud to be a Living Wage accredited Employer.
What You'll Do:
CoreWeave is building and operating some of the largest GPU infrastructure in the world. The Metal Net team owns the high-bandwidth GPU interconnect platforms that make large-scale AI and HPC workloads possible, including NVLink and NVSwitch-based systems. We deploy, operate, troubleshoot, and improve these platforms across our global data centre footprint to provide a powerful alternative to traditional hyperscalers.
About the role:
We are looking for an HPC Engineer to join our team to deploy, operate, and support NVLink/NVSwitch platforms across large data centre environments. This role is a strong fit for engineers who enjoy production troubleshooting, hardware-adjacent systems work, automation, observability, and learning specialized infrastructure deeply. You will be responsible for troubleshooting Linux, networking, hardware, firmware, performance, and stability issues in production, while building automation to improve runbooks, dashboards, alerts, and lifecycle workflows. Additionally, you will participate in rotating on-call shifts, lead incident responses, conduct root cause analyses, and collaborate cross-functionally across CoreWeave to ensure reliable workflows scale effectively as our global fleet grows.
Who You Are:
Strong Linux system administration and engineering troubleshooting skills.
Solid grasp of networking fundamentals and common diagnostic/troubleshooting tools.
Hands-on production debugging experience using logs, metrics, and command-line interfaces.
Technical experience troubleshooting server, network, GPU, or data centre hardware.
Practical scripting or automation experience using Python, Go, Bash, or similar languages.
Clear written and verbal communication, documentation skills, and readiness to participate in an on-call rotation.
High curiosity to deeply learn specialized GPU interconnect technologies such as NVLink, NVSwitch, and InfiniBand.
Preferred:
Experience with Ansible or other infrastructure-as-code and configuration automation tooling.
Kubernetes application development or live platform operations experience.
Familiarity with modern observability systems, including Grafana, Prometheus, PromQL, or similar stack components.
Experience managing large fleet operations across Linux systems, network devices, GPUs, or infrastructure components.
Deep understanding of InfiniBand, RDMA, HPC networking, or low-latency/high-bandwidth fabrics.
Experience with BMC, Redfish, IPMI, firmware lifecycle management, or hardware management APIs.
Exposure to NVLink, NVSwitch, NVIDIA GPU platforms, NVUE, SONiC, or specialized network operating systems.
Wondering if you're a good fit?
We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk.
You love to dive headfirst into production troubleshooting, hardware-adjacent systems work, and bringing robust automation to infrastructure at scale.
You're curious about specialized GPU interconnect technologies, high-bandwidth platforms, and continuous system optimization.
You're an expert in driving assigned work to completion with