Service Reliability Engineer
Location: Remote
Compensation: Salary
Reviewed: Mon, Aug 10, 2026
This job expires in: 17 days
Job Summary
Working remotely in a full-time capacity, the Service Reliability Engineer will manage extensive production GPU and Kubernetes environments, ensuring high availability and performance while delivering world-class support for on-prem and cloud products.
Key responsibilities
- Operate within a 24/7 support model, managing a flexible 4-day, 10-hour schedule to maintain global coverage
- Monitor and manage production environments, applying advanced tools for incident detection and resolution
- Develop predictive automated support routines and improve automation for incident management and system reliability
Required qualifications
- 8+ years of experience with large-scale production systems, including over 3 years in high-availability environments
- Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management
- Expert-level Linux system administration and proficiency in automation using Ansible and/or Python
- BS in Computer Science, Engineering, Physics, Mathematics, or equivalent experience
- Familiarity with observability and incident management tools such as Grafana, OpenTelemetry, and PagerDuty
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 45,615 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee