Inference
Genesis
| Company | Genesis |
| Category | Uncategorised |
| Location | Paris |
| Remote | Hybrid |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 30 Apr 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (ashby) |
Description
WHAT YOU’LL DO
- Build low-latency inference pipelines for on-device deployment, enabling real-time next-token and diffusion-based control loops in robotics
- Design and optimize distributed inference systems on GPU clusters, pushing throughput with large-batch serving and efficient resource utilization
- Implement efficient low-level code (CUDA, Triton, custom kernels) and integrate it seamlessly into high-level frameworks
- Optimize workloads for both throughput (batching, scheduling, quantization) and latency (caching, memory management, graph compilation)
- Develop monitoring and debugging tools to guarantee reliability, determinism, and rapid diagnosis of regressions across both stacks
WHAT YOU’LL BRING
- Deep experience in distributed systems, ML infrastructure, or high-performance serving (8+ years)
- Production-grade expertise in Python, with strong background in systems languages (C++/Rust/Go)
- Low-level performance mastery: CUDA, Triton, kernel optimization, quantization, memory and compute scheduling
- Proven track record scaling inference workloads in both throughput-oriented cluster environments and latency-critical on-device deployments
- System-level mindset with a history of tuning hardware–software interactions for maximum efficiency, throughput, and responsiveness
986,449 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →