Remote Jobs Sign In

Service Reliability Engineer

Location: Remote
Compensation: Salary
Reviewed: Mon, Aug 10, 2026
This job expires in: 17 days

Job Summary

Working remotely in a full-time capacity, the Service Reliability Engineer will manage extensive production GPU and Kubernetes environments, ensuring high availability and performance while delivering world-class support for on-prem and cloud products.

Key responsibilities
  • Operate within a 24/7 support model, managing a flexible 4-day, 10-hour schedule to maintain global coverage
  • Monitor and manage production environments, applying advanced tools for incident detection and resolution
  • Develop predictive automated support routines and improve automation for incident management and system reliability
Required qualifications
  • 8+ years of experience with large-scale production systems, including over 3 years in high-availability environments
  • Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management
  • Expert-level Linux system administration and proficiency in automation using Ansible and/or Python
  • BS in Computer Science, Engineering, Physics, Mathematics, or equivalent experience
  • Familiarity with observability and incident management tools such as Grafana, OpenTelemetry, and PagerDuty

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 45,615 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee