Staff Network Reliability Engineer
Location: Remote
Compensation: Salary
Reviewed: Tue, Jul 28, 2026
This job expires in: 30 days
Job Summary
Focused on maintaining cloud infrastructure health, the full-time Staff Network Reliability Engineer will oversee 24x7 operations across Skylo's hybrid production environment, ensuring reliability and observability while managing incidents and driving automation efforts remotely.
Key Responsibilities
- Own 24x7 cloud infrastructure health across hybrid environments, including GKE and on-premise Kubernetes clusters
- Lead incident diagnosis and escalation for Cloud Infra issues, serving as the L3 escalation authority
- Define and maintain SLOs for cloud infrastructure components, driving error budget management and toil reduction initiatives
Required Qualifications
- 8-10+ years of experience in infrastructure engineering or Site Reliability Engineering in a production 24x7 environment
- Deep expertise in Kubernetes, including multi-cluster operations and persistent storage management
- Experience with public cloud (GCP or AWS) and on-premise/private cloud infrastructure
- Proficient in observability tools such as Prometheus, Grafana, and OpenTelemetry
- Strong skills in database reliability, specifically with PostgreSQL and Redis operations
COMPLETE JOB DESCRIPTION
The job description is available to subscribers. Subscribe today to get the full benefits of a premium membership with Virtual Vocations. We offer the largest remote database online...