Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

AI Infrastructure Engineer

42dot
Company42dot
CategoryEngineering
LocationPangyo
RemoteHybrid
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted20 Jul 2026
Last verified30 Jul 2026
SourceEmployer career page (ashby)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
ABOUT THE TEAM & MISSION 42dot의 AI 인프라 엔지니어는 여러 데이터 센터에 걸쳐 있는 수천 개의 GPU를 관리하며, 이를 효율적으로 오케스트레이션하는 고성능 AI 인프라를 운영합니다. 세계 최고 수준의 컴퓨팅 환경을 유지하기 위해 확장성, 모니터링 및 운영 최적화 전반에 기여하게 됩니다. At 42dot, our AI Infrastructure Engineer manages the high-performance AI infrastructure orchestrating thousands of GPUs across multiple data centers. You will contribute to the scaling, monitoring, and operational optimization required to maintain a robust and world-class computing environment. RESPONSIBILITIES - Kubernetes 및 Slurm을 활용하여 여러 데이터 센터에 분산된 수천 개 규모의 대규모 GPU 클러스터 운영 및 유지 보수 - GPU 하드웨어 및 소프트웨어 스택 전반의 장애를 모니터링하고 진단하여 고가용성 유지 및 신속한 장애 복구 수행 - Python 또는 Shell을 활용한 자동화 도구 및 스크립트를 개발하여 반복적인 인프라 관리 업무를 효율화 - GPU 리소스 쿼터(Quota) 관리 및 ML 개발자를 위한 기술 지원을 통해 컴퓨팅 자원의 최적 활용 보장 - 대규모 자율주행 모델 학습을 위한 분산 학습 환경의 아키텍처 설계 및 성능 튜닝 참여 - Operate and maintain a large-scale GPU cluster consisting of thousands of GPUs across multiple data centers using Kubernetes and Slurm. - Monitor and diagnose failures across the GPU hardware and software stacks to ensure high availability and rapid recovery. - Develop automation tools and scripts using Python or Shell to streamline repetitive infrastructure management tasks and improve operational efficiency. - Manage GPU resource quotas and provide technical support to ML researchers to ensure optimal utilization of computing resources. - Participate in the architectural design and performance tuning of distributed training environments for large-scale autonomous driving models. QUALIFICATIONS - Linux 운영체제에 대한 깊은 이해 (커널 동작, 프로세스 관리, 시스템 보안 등) - Docker 및 Kubernetes 등 컨테이너 기반 기술 및 오케스트레이션 실무 경험 - TCP/IP, HTTP(S) 등 네트워크 기본 원리에 대한 이해 및 기초적인 네트워크 트러블슈팅 능력 - Python 또는 Shell을 활용하여 유지보수가 용이한 자동화/시스템 관리 스크립트 작성 역량 - 복잡하고 거대한 시스템에서 근본 원인을 찾아 해결하는 논리적인 문제 해결 능력 - 다양한 유관 부서 및 파트너와 원활하게 소통할 수 있는 커뮤니케이션 역량 - Strong proficiency in Linux operating systems, including a solid understanding of kernel operations, process management, and system security. - Practical experience with containerization technologies (Docker) and orchestration (Kubernetes), including building, managing, and troubleshooting containerized environments. - Solid understanding of networking fundamentals, including TCP/IP and HTTP(S), with the ability to perform basic network troubleshooting. - Ability to write clean and maintainable scripts in Python or Shell for automation and system administration. - Logical approach to problem-solving with the persistence to identify and resolve root causes in complex, large-scale systems. - Strong communication skills to effectively collaborate with cross-functional teams and external partners. PREFERRED QUALIFICATIONS - Prometheus, Grafana, Datadog 등을 활용한 대규모 클러스터의 관측성(Observability) 스택 구축 경험 - AWS, GCP 등 퍼블릭 클라우드 플랫폼 상의 인프라 구축 및 운영 경험 - 드라이버, CUDA, NCCL 등을 포함한 NVIDIA 가속 컴퓨팅 스택에 대한 지식 - ML 모델 학습 라이프사이클 및 PyTorch, TensorFl
HOUSE ADYou found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →
AI Infrastructure Engineer — 42dot · Job Opportunities API