Senior Site Reliability Engineer
Location: Remote
Compensation: Salary
Reviewed: Tue, Sep 29, 2026
This job expires in: 30 days
Job Summary
Seeking a full-time Senior Site Reliability Engineer to manage production reliability across a hybrid infrastructure platform, including cloud, colocation, and edge environments, while working remotely in the US and focusing on observability, automation, and incident response.
Key responsibilities:
- Own production reliability for customer-facing platforms and data services across various environments
- Build and maintain the observability layer, ensuring metrics and alerts are accessible and actionable
- Design automated recovery systems to enhance operational resilience and reduce downtime
Required qualifications:
- Bachelor's degree in computer science, software engineering, or a related field, or equivalent professional experience
- Minimum of 7 years of experience in Site Reliability Engineering or related roles, with 4 years specifically as a Site Reliability Engineer
- Deep hands-on experience operating native Kubernetes and optimizing cluster resources
- Proven ability to increase operational visibility through dashboards, metrics, and alerting
- Experience with incident response for customer-facing production systems and diagnosing complex production issues
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 38,810 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee