Site Reliability Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Fri, Sep 18, 2026
This job expires in: 30 days
Job Summary
Seeking an experienced Site Reliability Engineer to work remotely in a full-time capacity, this role focuses on building reliable, observable systems, managing incident response, and mentoring junior engineers while collaborating with cross-functional teams to enhance system performance and reliability.
Key Responsibilities
- Design observability systems and define SLOs, error budgets, and monitoring strategies aligned with business needs
- Own a 24/7 P0 on-call rotation with a 10-minute acknowledgment SLA, validating and escalating AI-generated incident reports
- Implement reliability improvements, capacity planning, and performance optimization while mentoring and onboarding additional SRE engineers
Required Qualifications
- 6+ years of hands-on experience in SRE, DevOps, or Platform Engineering at scale
- Strong proficiency in AWS services, including ALB, ECS/Fargate, and RDS Aurora
- Production experience in Python backend engineering, with skills in debugging and optimizing services
- Experience with incident response, on-call rotations, and building SRE programs from scratch
- Solid understanding of containerization technologies, including Kubernetes and Docker
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 41,590 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee