Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

AI Platform Support Engineer (APAC)

Lightning AI
CompanyLightning AI
CategoryUncategorised
LocationPhilippines
RemoteOn-site (inferred)
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted15 May 2026
Last verified30 Jul 2026
SourceEmployer career page (greenhouse)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction. Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in. We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute. The Way We Work The people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice: Move with Urgency: We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping. Take Ownership: We own outcomes, not just our individual work. We make decisions that move the company forward and follow through. Communicate Openly: We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together. Build Great Teams: We lead by example, empower others, and create healthy teams where people can do their best work. Raise the Bar: We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most. Think Long-Term: We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.   What We’re Looking For Lightning AI is looking to hire an AI Platform Support Engineer to join our APAC Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production environments. This role is not a ticket router or traditional support engineer. You are a technical partner to ML teams - helping diagnose failures, improve reliability, and guide customers through complex distributed systems problems.The problems range from Kubernetes scheduling and GPU orchestration to distributed PyTorch failures, inference latency, networking bottlenecks, storage performance, and platform reliability. You’ll gain exposure to a wide variety of real world AI workloads across industries and help shape the infrastructure powering the next generation of ML applications. This role is remote and open to candidates based in the Philippines. We are hiring for a Sunday-Wednesday shift schedule, with working hours from 7:00 AM to 5:00 PM local time (UTC+8). What You'll Do Work Directly With ML Engineers Partner directly with customer engineering teams running training and inference workloads in production Help customers diagnose and resolve complex distributed systems and ML infrastructure issues Act as a technical advisor during high impact incidents and platform degradation events Translate infrastructure level issues into actionable guidance for ML engineers Build credibility with customers through strong technical reasoning and clear communication Debug ML Infrastructure & Distributed Workloads Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems Troubleshoot PyTorch, CUDA, NCCL, and inference serving related issues Analyze logs, metrics, traces, and system behavior to isolate root causes Debug containerized workloads running across Kubernetes and bare metal GPU environments Support customers scaling workloads across multi node GPU systems Diagnose performance bottlenecks invo
HOUSE AD995,367 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →