Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Infrastructure Operations Engineer (APAC)

Lightning AI
CompanyLightning AI
CategoryEngineering
LocationSG
RemoteRemote
EmploymentNot stated
LevelNot stated
SalarySGD 165k–205k
Posted27 Jul 2026
Last verified12 Aug 2026
SourceThe employer's own careers page (company_site)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Lightning AI, the company behind PyTorch Lightning, is seeking an experienced Infrastructure Operations Engineer to join the APAC InfraOps team. This role is central to reliability, automation, and operational scale for GPU infrastructure, handling break/fix operations, incident response, customer provisioning, observability, and automation systems for large-scale infrastructure. What You'll Do • Design, build, and roll out new platforms and patterns to minimize incidents and enable customer-facing and internal features • Deploy updates and improvements to support Voltage Park's internal and end-customer use cases • Collaborate with Infrastructure Engineering, Network Operations, Customer Success, and Software Platform Development teams on troubleshooting and operational efficiency • Participate in on-call rotation (evenly distributed across team members in primary/secondary pattern) to respond to incidents and manage GPU infrastructure reliability • Build automation systems that reduce manual toil and improve operational efficiency across large-scale GPU environments and bare metal infrastructure What You Need • 8+ years working with Linux as a server/hosting platform • 5+ years experience with AWS • 2+ years experience with Kubernetes and strong container fundamentals • 2+ years experience with Terraform and Ansible • 2+ years with network attached storage management (NFS, Ceph, or equivalent protocols) • Experience with monitoring systems (Prometheus, ELK stack) • Familiarity with GitOps workflow • Software development experience using Python, Go, Bash, or other languages for automation and system integration • Deep networking fundamentals including datacenter-level networking knowledge • Experience building and delivering complex systems with ability to navigate tradeoffs between design, risk, cost, and outcomes • Strong written and oral communication skills Nice to Have • Experience with bare metal hardware troubleshooting and provisioning, particularly Dell hardware • Experience with GPU servers in bare metal or virtualized form • Deep experience with network switches, routers, and firewalls (SONiC, Palo Alto, Juniper Networks) • Experience with VAST storage systems $165,000–$205,000 SGD annually. Total rewards include discretionary bonus, meaningful equity (RSUs), comprehensive health coverage (medical, dental, vision), 401(k) matching (U.S.) and pension contributions (U.K.), unlimited PTO, company holidays, paid parental and family leave, annual learning and development allowance, wellness and work-from-home stipends, four weeks paid sabbatical after four years of service, flexible schedules, and complimentary meals at office hubs.