Staff Site Reliability Engineer
Location: Remote
Compensation: Salary
Reviewed: Fri, Sep 18, 2026
This job expires in: 30 days
Job Summary
As the first dedicated Staff Site Reliability Engineer in a fully remote capacity, the successful candidate will manage reliability metrics, enhance incident response processes, and foster a reliability-focused mindset across engineering teams, ensuring operational excellence and measurable reliability standards.
Key responsibilities
- Define and implement SLIs and SLOs for critical request paths, ensuring visibility and accountability among teams
- Strengthen the incident lifecycle by improving detection, response, and postmortem processes while driving alert quality and escalation design
- Coach teams on reliability practices, embedding a culture of operational excellence and deliberate failure testing
Required qualifications
- 10+ years of engineering experience, including 3+ years as a Site Reliability Engineer or in a similar reliability-focused role
- Proven expertise in SLI/SLO design and error budgets, with a track record of team adoption
- Strong incident leadership experience, having managed high-severity incidents and improved organizational learning
- Hands-on experience with distributed systems and proficiency in Kubernetes, AWS, and modern observability tools
- Ability to read and write production code (Go, TypeScript, or similar) and familiarity with infrastructure as code
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 41,590 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee