Sr. Software Engineer-AI Reliability
MixMode
| Company | MixMode |
| Category | Engineering |
| Location | Remote |
| Remote | Remote |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 3 Mar 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (greenhouse) |
Description
MixMode is a leading provider of AI-powered cybersecurity solutions at scale, pioneering a patented third-wave, context-aware AI approach that automatically learns and adapts to dynamic environments. The MixMode platform delivers self-supervised, real-time threat detection for known and unknown threats across cloud, hybrid, and on-premises environments. Large organizations with big data workloads – including those in enterprise, critical infrastructure, US Department of War and US Intelligence Community – trust MixMode to defend their most important assets. Backed by PSG and Entrada Ventures, MixMode is headquartered in Santa Barbara, California. Learn more at www.mixmode.ai . Job Summary:
We are hiring a Senior Software Engineer, based in the United States, to enhance the reliability, performance, and scalability of our production AI systems. We value clear thinking, incremental improvement, and engineers with real production incident experience. This role focuses on understanding, refining, and strengthening existing distributed services across application, database, and container orchestration layers. You will collaborate with ML researchers to make our systems more robust, maintainable, flexible, and scalable. This is a remote position but requires the ability to travel to the MixMode headquarters in Santa Barbara for in person meetings a few times a year.
What you’ll be doing (responsibilities):
Own the reliability, performance, and operational health of production AI services
Refactor and harden existing systems to improve resilience, clarity, and maintainability
Diagnose and resolve issues across distributed services, data pipelines, and storage layers
Design and implement monitoring, alerting, and debugging tools for high-availability systems
Partner with researchers and engineers to productionize predictive systems at scale
Establish best practices for testing, deployment, capacity planning, and incident response
Contribute to incident response and postmortems, driving continuous improvement
What you’ll need to bring (qualifications):
Ability to travel to our office in Santa Barbara, CA, a few times per year
7+ years of professional software engineering experience
Strong proficiency in Python and at least one JVM language (Java, Scala, Kotlin)
Proven experience designing, building, and operating distributed systems in production
Strong understanding of service architecture, concurrency, resource management, and distributed failure modes
Experience operating Kubernetes deployments
Strong experience with relational databases, including query performance analysis, indexing, and connection management
Demonstrated ability to diagnose and resolve performance, scalability, and reliability issues across system layers
Experience implementing automated testing and production observability (logging, metrics, tracing)
Experience collaborating with ML or data science teams (deep ML expertise is not required)
Ability to improve system architecture and engineering practices through design, code review, and mentorship
Our Interview Process:
Our interview process focuses on real-world production experience and practical systems thinking. We assess how you reason about distributed systems, refactor existing code, and operate under real constraints—not abstract puzzles. We support the use of AI tools in our development, but we want to understand your capabilities first. No AI tools will be allowed for remote interviews early in the process.
The process includes:
Conversations about systems you’ve personally owned, improved, and operated in production
A live refactoring and testing exercise in your choice of Java, Kotlin, or Scala, centered on improving existing code without changing behavior
A distributed systems discussion covering performance, state management, failure modes, and debugging under load
An ML production discus
1,014,484 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →