Remote Jobs Sign In

Staff Site Reliability Engineer

Location: Remote
Compensation: Salary
Reviewed: Fri, Sep 18, 2026
This job expires in: 30 days

Job Summary

Passionate about building resilient systems at scale, the full-time Staff Site Reliability Engineer will ensure the reliability, scalability, and performance of infrastructure by implementing automation, leading incident response, and mentoring the engineering team, all while working remotely.

Key responsibilities
  • Architect and implement comprehensive monitoring, logging, and tracing solutions for real-time system health visibility
  • Define and drive reliability standards by collaborating with product and engineering teams to establish Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Lead incident management and response efforts, guiding teams during high-impact incidents and conducting post-mortems to implement preventative measures
Required qualifications
  • 8-10 years of experience in Site Reliability Engineering or similar roles
  • Strong programming skills in Python or Go, with a focus on writing high-quality, well-tested code
  • Deep understanding of distributed systems and experience with container orchestration platforms, particularly Kubernetes
  • Proven track record in designing and implementing sophisticated monitoring and observability solutions
  • Strong incident management skills with extensive experience leading incident response for complex systems

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 41,725 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee