Site Reliability Engineer (USA Only - 100% Remote)
Close
| Company | Close |
| Category | Engineering |
| Location | USA - Remote |
| Remote | Remote |
| Employment | Not stated |
| Level | Not stated |
| Salary | USD 140k–210k |
| Posted | 30 Mar 2026 |
| Last verified | 8 Aug 2026 |
| Source | Employer ATS (ashby) |
Description
ABOUT US
Since 2013, we’ve been building a CRM that gets out of your way and helps your team sell more, faster. Now we’re building AI into every part of it, so Close does the busywork and your team does the selling. No manual data entry, no 10-click workflows. Just communication-first, AI-powered sales software designed to help you succeed and scale.
We're bootstrapped and profitable which means we answer to our customers and play by our rules. We're proud of our 120-person, 100% remote team, focused on building Close so that no small, scaling business fails because it can't figure out sales.
We love open sourcing our code and ideas on our GitHub https://github.com/closeio and on The Making of Close https://making.close.com/, our behind-the-scenes Product & Engineering blog. Check out our open source projects like close-mongo-ops-manager https://github.com/closeio/close-mongo-ops-manager, SocketShark https://github.com/closeio/socketshark, TaskTiger https://github.com/closeio/tasktiger, LimitLion https://github.com/closeio/limitlion and ciso8601 https://github.com/closeio/ciso8601.
AI is both how we build and what we ship, and that's reshaped what engineering looks like. This is a transformation we're embracing and find deeply exciting.
OUR PLATFORM
We run the platform every other system at Close depends on: multi-terabyte MongoDB, PostgreSQL, and Elasticsearch clusters, multiple Kubernetes clusters running tens of thousands of pods, and telemetry on Grafana's LGTM stack and ClickHouse processing over 130 TB a month. We run AI in production for 11,000 paying customers, and the system underneath hasn't needed scheduled downtime in 4 years.
The underlying infrastructure runs on AWS using a combination of managed services like EKS, MSK, RDS and ElastiCache and self-managed services running on EC2 instances. We have CI/CD pipelines that build Docker images, run automated tests and deploy to our clusters. We also use these images in our local development environment allowing coding locally against all of our services.
We have a well-documented public API https://developer.close.com/ that is consumed by our front-end TypeScript app as well as numerous integrations. Our infrastructure is heavily automated using Terraform, Ansible and other AWS tools.
YOU WILL
- Automate our databases' full lifecycle. Provisioning, scaling, failover, retirement — multi-terabyte MongoDB, PostgreSQL, and Elasticsearch, handled by the platform instead of by hand.
- Drive static credentials out of the system wherever they live. Move us toward short-lived, identity-based auth across services and infrastructure, shrinking the blast radius every quarter as Close moves upmarket.
- Push downtime and disruption to new lows. Maintenance, deploys, and disaster recovery that customers never feel — building on a system that hasn't needed scheduled downtime in 4 years.
- Harden our multi-region disaster recovery. Make failover faster, more automatic, and more trustworthy, so a regional outage is a non-event.
- Own the telemetry and CI/CD backbone the whole company runs on. The Grafana LGTM + ClickHouse pipeline processing 130 TB a month, and the GitHub Actions + ArgoCD path that goes merged-to-production-to-rolled-back in 10 minutes.
YOU ARE
- A rock in the storm. With hard-won expertise gained through battles won and lost, you build robust systems from quality components fit to underpin mission-critical applications. You value simplicity over familiarity and resilience over speed, and you take pride in composable, maintainable tools.
- Fluent across the infrastructure stack. You've worked with a wide range of tools and systems: CI/CD (GitHub Actions, ArgoCD), configuration management (Ansible, Terraform), databases (Elasticsearch, MongoDB, PostgreSQL, ClickHouse), cloud (Kubernetes, AWS), and telemetry (Loki, Tempo, Grafana, Mimir/Prometheus, OTEL). You pick the right tool for the problem rather than retreating to what you k