Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Senior AI Infrastructure Support Engineer - APAC (GPUs)

Nscale
CompanyNscale
CategoryEngineering
LocationSingapore
RemoteOn-site (inferred)
EmploymentNot stated
LevelSenior
SalaryNot stated by the employer
Posted22 Jun 2026
Last verified12 Aug 2026
SourceEmployer ATS (greenhouse)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers.  Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility. At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability and rapid response to customer tickets.  We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future. Senior AI Infrastructure Support Engineer - APAC What You'll be Doing Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes Diagnose and remediate GPU node faults across the full stack — driver, firmware, and hardware layers — from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics. Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage–network interactions across clients, mounts, VIPs, and routing Run structured, hypothesis-driven investigations; conduct root cause analysis for major incidents and drive long-term fixes to completion Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents; identify recurring patterns and convert them into problem records and automation Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover Design and implement automation scripts and small tools to reduce toil and human intervention Act as a key escalation point for the Support Organisation; taking ownership of strategic decisions where results matter Mentor and upskill mid-level engineers; contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo; commit wholly to decisions once in motion Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise About You 6+ years in infrastructure, operations, or support engineering in production environments; 2–3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity Able to explain complex technical detail clearly, specifically, and concisely — in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch GPU platforms (NVIDIA; AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM,