Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

HPC Network Engineer

Fuse Energy
CompanyFuse Energy
CategoryEngineering
LocationLondon
RemoteOn-site (inferred)
EmploymentFull-time
LevelNot stated
SalaryNot stated by the employer
Posted23 Jul 2026
Last verified2 Aug 2026
SourceEmployer ATS (workable)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast. We're combining first-principles thinking with cutting-edge technology to build a radically better energy system. We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more. The Opportunity You'll design, deploy, and operate the network fabric for our multi-tenant AI cluster. This covers the full stack: the high-performance compute and storage fabrics carrying RDMA traffic between GPUs, the tenant-facing and management networks, fire walling and tenant isolation, and the out-of-band infrastructure that keeps it all recoverable. Beyond the data centre, you'll own the office network and act as the networking authority for the company, raising the bar for everyone by sharing what you know. You'll own the fabric from architecture through day-2 operations. Responsibilities Design and operate lossless, RDMA-capable fabrics (e.g. RoCEv2, InfiniBand) for GPU compute and storage traffic, including QoS, congestion control, and buffer tuning at scale Build and manage leaf-spine data centre fabrics, with routed underlay and overlay design (e.g. BGP, EVPN/VXLAN) Implement and maintain per-tenant network isolation across compute, storage, and management planes. Automate network provisioning, configuration, and validation, treating switch config as code (e.g. Ansible, Python, NetBox as source of truth), deployed through CI Build telemetry and observability for the fabric: flow-level and buffer-level visibility, dashboards, and alerting that catches congestion and link degradation before tenants do (e.g. Prometheus/Grafana/Datadog, streaming telemetry) Troubleshoot performance issues end to end, from optics and cabling through switch buffers to NIC/DPU configuration and collective-communication behaviour on the hosts Operate the out-of-band management network, console access, and remote recovery paths Support tenant onboarding: segmentation and addressing, bandwidth and isolation guarantees, and capacity planning as the cluster scales Write clear design documentation capturing decisions, rationale, and rejected alternatives Own and maintain the office network: wired and wireless infrastructure, firewalling, VPN/remote access, and connectivity between the office and data centre environments Upskill colleagues on networking: share knowledge through documentation, run-throughs, and pairing so the wider team can operate and troubleshoot the fabric confidently