Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Staff Network Reliability Engineer, Cloud Operations

Skylo
CompanySkylo
CategoryEngineering
LocationMountain View
RemoteHybrid
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted23 Jul 2026
Last verified30 Jul 2026
SourceEmployer career page (ashby)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
ABOUT SKYLO Skylo is a global Non-Terrestrial Network (NTN) service provider based in Mountain View, CA, offering a service that allows smartphone and IoT cellular devices to connect directly over existing satellites. Skylo's direct-to-device service is live on millions of activated devices across five continents, with more than 60 million square kilometers of coverage, in partnership with multiple satellite operators, mobile network operators (MNOs), Tier-1 chipset makers, and OEMs. Devices connected over satellite are managed and served by Skylo's commercial NTN vRAN — a 3GPP standards-based, cloud-native base station and core. Skylo provides an anywhere, anytime connectivity solution that seamlessly roams between terrestrial and satellite networks. Our focus is on enabling connected services across three main verticals: mass-market consumer devices, automotive, and industrial IoT. HOW YOU WILL IMPACT SKYLO As a Staff Network Reliability Engineer, Cloud Operations, in the Global Product Support & Customer Success organization, you are the Cloud Infrastructure domain authority within Skylo's production NTN network. Everything runs on the infrastructure you keep healthy — RAN NFs, Core NFs, OSS, BSS, and the observability pipeline itself. When a GKE node fails, when ArgoCD drifts, when a Persistent Volume Claim goes unavailable, when a PostgreSQL replica falls behind, when Prometheus WAL corrupts — you own the response. You operate across Skylo's full hybrid cloud estate: GCP public cloud (GKE clusters, Pub/Sub pipelines, Cloud SQL) and on-premise private cloud infrastructure (bare-metal Kubernetes, hyperconverged compute, software-defined storage). You own 24x7 platform health, the observability pipeline (Prometheus, VictoriaMetrics, Grafana, OpenTelemetry), persistent storage operations (PostgreSQL, Redis), and the operational interface with Network Implementation for all GitOps-driven infrastructure changes. At Staff NRE level you go beyond keeping the lights on. You define and maintain SLOs tied to network SLA commitments, own error budget tracking, drive toil reduction, and partner with Ops Platform Engineering to automate the infrastructure remediation loop. You are the last technical stop before a Cloud Infrastructure problem becomes an engineering escalation — and you are a force multiplier, mentoring Senior NREs and contributing the operational requirements that shape Skylo's platform roadmap. KEY RESPONSIBILITIES CLOUD INFRASTRUCTURE OPERATIONS & HEALTH OWNERSHIP - Own 24x7 cloud infrastructure health across Skylo's hybrid production environment: GKE cluster node status, namespace and pod health, Persistent Volume Claim availability, network policies, and multi-cluster federation across Skylo's GCP footprint. - Own on-premise Kubernetes cluster health: bare-metal node availability, container runtime stability, CNI networking, persistent storage arrays (Ceph/Rook or equivalent), and hyperconverged compute platform operations (Harvester, KubeVirt, or KVM). - Monitor and triage infrastructure alarms using OSS dashboards, Grafana/VictoriaMetrics telemetry, GCP Cloud Monitoring, and Loki log correlation — distinguish transient platform events from systemic infrastructure degradation. - Execute and own Cloud Infra runbooks for P2–P4 fault categories: GKE node recovery, pod eviction and rescheduling, PVC repair, database failover execution, Prometheus WAL corruption recovery, ArgoCD drift remediation, and certificate rotation — without requiring engineering involvement for covered fault classes. - Own BSS-IIS GKE cluster monitoring and infrastructure health; maintain runbooks that reflect current cluster topology after every infrastructure change. OBSERVABILITY PIPELINE & DATA PLATFORM OPERATIONS - Own the observability pipeline end-to-end: Prometheus scrape target integrity, VictoriaMetrics retention and query performance, Grafana dashboard coverage and accuracy, Ope
HOUSE ADYour CV gets thirty seconds.CV writing and honest review. English & Greek.kaeros.app →