Remote Jobs Sign In

Principal Site Reliability Engineer

Location: Remote
Compensation: To Be Discussed
Reviewed: Thu, Sep 17, 2026
This job expires in: 30 days

Job Summary

Leading the reliability and performance of production systems, the full-time remote Principal Site Reliability Engineer will manage incident response, drive infrastructure automation, and enhance operational governance while collaborating with cross-functional teams.

Key responsibilities
  • Oversee day-to-day production support and incident management, including escalation and stakeholder communication
  • Define and maintain SLIs and SLOs for critical services to guide monitoring and reliability efforts
  • Lead infrastructure automation initiatives using Terraform, ensuring best practices in CI/CD pipelines and operational governance
Required qualifications
  • 10+ years of experience in Site Reliability Engineering, DevOps, or infrastructure engineering
  • Expertise in Terraform or comparable Infrastructure as Code (IaC) at scale
  • Strong grounding in SRE principles, including incident management and blameless postmortems
  • Hands-on experience with cloud platforms (AWS, Azure, or GCP) and container orchestration (Docker/Kubernetes)
  • Proficiency in at least one scripting or programming language (e.g., Python, Go, Bash) for automation

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 41,606 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee