Staff Site Reliability Engineer
Location: Remote
Compensation: Salary
Reviewed: Wed, Aug 19, 2026
This job expires in: 30 days
Job Summary
Seeking a remote Staff Site Reliability Engineer, the successful candidate will own the reliability strategy for cloud infrastructure, lead incident response programs, and architect observability platforms while working closely with AI/ML workloads.
Key responsibilities
- Own the reliability strategy for cloud environments, including designing SLO frameworks and leading technical decision-making
- Lead the incident response program, managing complex escalations and driving root cause analysis
- Architect the monitoring and observability platform to proactively detect and resolve issues
Required qualifications
- 7+ years of experience operating production cloud infrastructure in SRE, DevOps, or platform engineering roles
- Deep expertise with Kubernetes and Terraform, particularly in AWS environments
- Experience designing reliability practices, including SLO frameworks and incident response programs
- Strong programming skills in Python or Go for infrastructure automation
- Mentorship experience with the ability to set technical direction for teams
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 47,805 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee