Principal Site Reliability Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Thu, Sep 17, 2026
This job expires in: 30 days
Job Summary
Leading the reliability and performance of production systems, the full-time remote Principal Site Reliability Engineer will manage incident response, drive infrastructure automation, and enhance operational governance while collaborating with cross-functional teams.
Key responsibilities
- Oversee day-to-day production support and incident management, including escalation and stakeholder communication
- Define and maintain SLIs and SLOs for critical services to guide monitoring and reliability efforts
- Lead infrastructure automation initiatives using Terraform, ensuring best practices in CI/CD pipelines and operational governance
Required qualifications
- 10+ years of experience in Site Reliability Engineering, DevOps, or infrastructure engineering
- Expertise in Terraform or comparable Infrastructure as Code (IaC) at scale
- Strong grounding in SRE principles, including incident management and blameless postmortems
- Hands-on experience with cloud platforms (AWS, Azure, or GCP) and container orchestration (Docker/Kubernetes)
- Proficiency in at least one scripting or programming language (e.g., Python, Go, Bash) for automation
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 41,606 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee