Infrastructure Support Engineer
Nscale
| Company | Nscale |
| Category | Engineering |
| Location | Houston; New York; San Francisco; Seattle |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 4 Jun 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (greenhouse) |
Description
About Nscale
Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world.
At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.
About the Role
Infrastructure Support Engineers are the delivery engine of Infrastructure Support (L2/L3), handling the day-to-day health of Nscale's GPU fleets — tickets, alerts, hardware faults, and customer issues — across GPU nodes, high-performance networks, Linux, and data centre operations. This is a hands-on technical role: you'll come in with a strong technical base and working knowledge of GPU infrastructure, and grow toward Senior through exposure to some of the most advanced AI infrastructure in the world.
You will:
Own your tickets and tasks end-to-end, escalating early and appropriately when issues exceed your scope — with clean, evidence-rich handovers.
Communicate technical detail clearly, specifically, and concisely — in tickets, to customers, and to colleagues. We treat communication quality as a core engineering skill, not a soft skill.
Grasp new technical concepts and problems quickly; stay curious and know what questions to ask to get up to speed fast.
Bring discipline and organisation: accurate records, structured troubleshooting, reliable follow-through.
Seek feedback and invest in learning — this role is a deliberate pathway to Senior.
Experience required: 3-4+ years in infrastructure support or support engineering roles, including deep support/service desk experience in structured, customer-facing environments, with working knowledge of GPU infrastructure and hands-on hardware troubleshooting.
What You’ll Be Doing
Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately. Collaborate with Engineering, with guidance, when incidents or changes require it.
Perform GPU node triage and hardware troubleshooting: interpret nvidia-smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA.
Run fabric and link diagnostics following established runbooks (mlxlink or equivalent), capture evidence accurately, and escalate with a handover that lets the next engineer continue without starting from scratch.
Assist with storage and data-path investigations (mounts, connectivity, client-side symptoms) on high-performance platforms, gathering evidence for Senior or Engineering-led diagnosis.
Follow established runbooks to resolve common issues; propose improvements and contribute incremental fixes with review.
Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels.
Participate in monitoring, troubleshooting, and triage. Capture logs and facts to enable efficient handover.
Participate in changes under peer review, learning risk assessment and backout practices in live customer environments.
Help maintain source-of-truth accuracy across DCIM, inventory, and asset r
991,236 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →