Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Infrastructure Engineer

AION
CompanyAION
CategoryEngineering
LocationBengaluru
RemoteOn-site (inferred)
EmploymentFull-time
LevelNot stated
SalaryNot stated by the employer
Posted14 Jul 2026
Last verified3 Aug 2026
SourceEmployer career page (workable)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
About aion Aion is the enterprise AI platform, a full-stack solution for building, fine-tuning, and deploying AI at scale. Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform. We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters. We’re a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we’re building the next generation of enterprise AI and we’re looking for exceptional people to help us scale. Who You Are You are a visionary infrastructure architect passionate about democratizing AI compute at global scale. You thrive on solving complex technical challenges that create elegant, accessible systems from intricate infrastructure. With deep expertise in secure multi-tenancy environments, you understand how to design and implement comprehensive isolation guarantees across hardware, network, and storage layers for both VM and container workloads. You're excited to join an ambitious AI infrastructure startup at the ground floor, where your work will directly unlock siloed compute resources and remove barriers limiting AI advancement. You have the technical depth to architect platform systems that seamlessly connect compute providers with AI engineers while maintaining robust security foundations that scale to serve diverse client requirements and compliance needs. You're motivated by the opportunity to build something transformative—creating the infrastructure that will make high-performance compute more accessible, affordable, and user-friendly for the next generation of AI innovation. What You'll Do Observability Systems: Build and deploy comprehensive monitoring for GPU infrastructure using DCGM, NVML, and custom exporters; design metrics collection pipelines that track GPU health, utilization, thermal management, and performance across heterogeneous providers LGTM Stack Architecture & Operations: Deploy and manage production-scale Loki, Grafana, Tempo, and Mimir alongside Prometheus and Thanos/VictoriaMetrics; design retention strategies, aggregation rules, and query patterns that scale to thousands of GPUs Custom Exporter Development: Write Prometheus exporters in Go or Python for GPU metrics, platform services, and infrastructure components; implement proper metric naming, labeling strategies, and follow OpenMetrics standards Kubernetes Controller Development: Build custom controllers and operators for GPU workload management, scheduling, and resource allocation; instrument controllers with comprehensive metrics and tracing for observability Training & Inference Monitoring: Design and implement specialized observability for AI training workloads (GPU efficiency, distributed training performance, resource utilization) and inference services (latency percentiles, throughput, cost analytics) SLURM Integration & Monitoring: Deploy and manage SLURM clusters for HPC workloads, build observability for batch jobs, create bridges between SLURM and Kubernetes, and design unified monitoring across orchestrators Systemd Service Development: Write and deploy systemd services for bare-metal GPU nodes including monitoring agents, metric collectors, and platform daemons; implement proper logging and error handling GitOps Platform Engineering: Manage infrastructure using ArgoCD and GitOps workflows, create observable platform abstractions that teams consume declaratively, build self-service capabilities with integrated monitoring Alerting & SRE: Design intelligent, actionable alerting systems for GPU failures, thermal throttling, performance degradation, and workload anomalies; define platform SLOs and implement comprehensive monitoring to track
HOUSE AD1,970,675 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →