AI Infrastructure Engineer
42dot
| Company | 42dot |
| Category | Engineering |
| Location | Pangyo |
| Remote | Hybrid |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 20 Jul 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (ashby) |
Description
ABOUT THE TEAM & MISSION
42dot의 AI 인프라 엔지니어는 여러 데이터 센터에 걸쳐 있는 수천 개의 GPU를 관리하며, 이를 효율적으로 오케스트레이션하는 고성능 AI 인프라를 운영합니다. 세계 최고 수준의 컴퓨팅 환경을 유지하기 위해 확장성, 모니터링 및 운영 최적화 전반에 기여하게 됩니다.
At 42dot, our AI Infrastructure Engineer manages the high-performance AI infrastructure orchestrating thousands of GPUs across multiple data centers. You will contribute to the scaling, monitoring, and operational optimization required to maintain a robust and world-class computing environment.
RESPONSIBILITIES
- Kubernetes 및 Slurm을 활용하여 여러 데이터 센터에 분산된 수천 개 규모의 대규모 GPU 클러스터 운영 및 유지 보수
- GPU 하드웨어 및 소프트웨어 스택 전반의 장애를 모니터링하고 진단하여 고가용성 유지 및 신속한 장애 복구 수행
- Python 또는 Shell을 활용한 자동화 도구 및 스크립트를 개발하여 반복적인 인프라 관리 업무를 효율화
- GPU 리소스 쿼터(Quota) 관리 및 ML 개발자를 위한 기술 지원을 통해 컴퓨팅 자원의 최적 활용 보장
- 대규모 자율주행 모델 학습을 위한 분산 학습 환경의 아키텍처 설계 및 성능 튜닝 참여
- Operate and maintain a large-scale GPU cluster consisting of thousands of GPUs across multiple data centers using Kubernetes and Slurm.
- Monitor and diagnose failures across the GPU hardware and software stacks to ensure high availability and rapid recovery.
- Develop automation tools and scripts using Python or Shell to streamline repetitive infrastructure management tasks and improve operational efficiency.
- Manage GPU resource quotas and provide technical support to ML researchers to ensure optimal utilization of computing resources.
- Participate in the architectural design and performance tuning of distributed training environments for large-scale autonomous driving models.
QUALIFICATIONS
- Linux 운영체제에 대한 깊은 이해 (커널 동작, 프로세스 관리, 시스템 보안 등)
- Docker 및 Kubernetes 등 컨테이너 기반 기술 및 오케스트레이션 실무 경험
- TCP/IP, HTTP(S) 등 네트워크 기본 원리에 대한 이해 및 기초적인 네트워크 트러블슈팅 능력
- Python 또는 Shell을 활용하여 유지보수가 용이한 자동화/시스템 관리 스크립트 작성 역량
- 복잡하고 거대한 시스템에서 근본 원인을 찾아 해결하는 논리적인 문제 해결 능력
- 다양한 유관 부서 및 파트너와 원활하게 소통할 수 있는 커뮤니케이션 역량
- Strong proficiency in Linux operating systems, including a solid understanding of kernel operations, process management, and system security.
- Practical experience with containerization technologies (Docker) and orchestration (Kubernetes), including building, managing, and troubleshooting containerized environments.
- Solid understanding of networking fundamentals, including TCP/IP and HTTP(S), with the ability to perform basic network troubleshooting.
- Ability to write clean and maintainable scripts in Python or Shell for automation and system administration.
- Logical approach to problem-solving with the persistence to identify and resolve root causes in complex, large-scale systems.
- Strong communication skills to effectively collaborate with cross-functional teams and external partners.
PREFERRED QUALIFICATIONS
- Prometheus, Grafana, Datadog 등을 활용한 대규모 클러스터의 관측성(Observability) 스택 구축 경험
- AWS, GCP 등 퍼블릭 클라우드 플랫폼 상의 인프라 구축 및 운영 경험
- 드라이버, CUDA, NCCL 등을 포함한 NVIDIA 가속 컴퓨팅 스택에 대한 지식
- ML 모델 학습 라이프사이클 및 PyTorch, TensorFl
You found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →