Senior Site Reliability Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Thu, Oct 01, 2026
This job expires in: 30 days
Job Summary
Owning the reliability of production systems, the full-time Senior Site Reliability Engineer will define SLOs, lead incident response, and build tooling for safe shipping and experimentation, while working remotely with cross-functional teams to ensure infrastructure security and compliance across AWS and Azure.
Key responsibilities
- Define SLOs and SLIs with stakeholders, managing incident response and leading blameless postmortems
- Design, build, and maintain core infrastructure to enable rapid and safe product changes
- Collaborate with engineering teams to establish delivery standards, including CI/CD and safe deployment practices
Required qualifications
- 8+ years in site reliability, DevOps, or infrastructure engineering with end-to-end ownership of production systems
- Experience managing AWS and Azure resources in accordance with the Well-Architected Framework
- Proficiency in deploying Docker-based software and managing traditional data centers
- Deep expertise in SQL and relational database administration, including performance tuning
- Experience with infrastructure-as-code standards, particularly using Terraform
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 41,336 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee