Data Scientist, AI/ML
Gremlin
| Company | Gremlin |
| Category | Data & Analytics |
| Location | Remote |
| Remote | Remote |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 14 Jul 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (greenhouse) |
Description
Data Scientist, AI/ML
Job Description:
Today’s complex, fast-paced systems have become a minefield of reliability risks, any of which could cause an outage that costs millions and destroys customer confidence. That’s why high-availability teams use Gremlin to find and fix reliability risks before they become incidents.
Gremlin Reliability Platform helps software teams proactively monitor and test their systems for common reliability risks, build and enforce reliability standards, and automate their reliability practices organization-wide. As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the world’s largest organizations where high availability is non-negotiable.
About the Role of the Data Scientist, AI/ML
As a Data Scientist, AI/ML at Gremlin, you will have the opportunity to improve the reliability of the internet at large by turning millions of chaos engineering experiments into automated failure analysis and remediation. You will be able to leverage your applied machine learning experience to inform product direction as well as solve complex technical problems that directly impact our customers (which range from the Fortune 500 to smaller organizations). You will work closely with a small, talented engineering team focused on quality, delivery, and predictability with an emphasis on providing our customers a great user experience.
In this role, you’ll get to:
Analyze Gremlin’s proprietary dataset of millions of chaos engineering experiments to identify failure patterns, root causes, and resilience signals across complex distributed systems
Pretraining and fine-tuning machine learning models that automatically detect, classify, and explain failures observed during chaos experiments
Build intelligent systems that deliver automated remediation recommendations, and eventually orchestration, by learning from historical experiment outcomes and system behavior
Develop scalable data pipelines and feature stores to process, enrich, and serve large volumes of experiment data for both model training and real-time inference
Collaborate closely with platform engineers and SREs to integrate AI-driven failure analysis and remediation capabilities directly into Gremlin’s core product
Apply advanced techniques, including causal inference, graph ML, time-series modeling, and reinforcement learning, to continuously improve the accuracy and actionability of automated failure analysis
Translate insights from millions of chaos experiments into AI-powered features that help customers automatically understand blast radius, pinpoint root causes, and accelerate recovery
Research and productionize novel ML approaches, including causal AI and agentic systems, that turn raw chaos experiment data into automated, reliable remediation strategies
We’ll expect you to have:
Experience as a self-driven and collaborative problem solver with strong communication skills
5+ years professional experience building and productionizing machine learning, ideally for distributed systems, infrastructure, or DevOps and SRE use cases with more overall years of experience in software development.
Hands-on experience with techniques such as causal inference, graph ML, time-series modeling, or reinforcement learning
Experience building data pipelines and feature stores that support both offline training and real-time inference
Experience with agile development environments and practices
Strong advocate and practitioner of rigorous experimentation, model evaluation, and engineering best practices
Comfort partnering with platform engineers and SREs to turn research into shipped product features
Strong at breaking down ambiguous problems into concrete actions and milestones
Bonus Experience:
Experience with chaos engineering, site reliability engineering, or distributed systems
Background in agentic AI systems or large-scale causal inference
995,367 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →