Senior Site Reliability Engineer
Location: Remote
Compensation: Salary
Reviewed: Tue, Aug 04, 2026
This job expires in: 15 days
Job Summary
Leading cross-system triage and architectural remediation for complex production incidents, the full-time remote Senior Site Reliability Engineer will manage cloud infrastructure design, improve service reliability, and eliminate operational toil through automation.
Key responsibilities
- Serve as the senior escalation point for complex production incidents, leading permanent architectural remediation
- Lead platform-level architecture reviews to ensure cloud infrastructure meets reliability, scalability, and security standards
- Identify systemic failure patterns and implement architectural changes for lasting platform improvements
Required qualifications
- 8+ years in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering with ownership of complex production platforms
- Expert-level experience in designing and operating cloud infrastructure in Microsoft Azure, with knowledge of AWS
- Deep expertise in running Kubernetes in production, including cluster management and workload operations
- Advanced skills in Terraform and Infrastructure-as-Code (IaC), including module design and state management
- Strong scripting skills in Bash for automation and operational tooling development
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 46,254 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee