Remote Jobs Sign In

Principal Site Reliability Engineer

Location: Remote
Compensation: To Be Discussed
Reviewed: Fri, Aug 28, 2026
This job expires in: 30 days

Job Summary

To enhance the reliability and performance of production systems, the full-time Principal Site Reliability Engineer will lead incident management, implement SRE practices, and drive infrastructure automation in a fully remote environment.

Key responsibilities:
  • Manage day-to-day production support and incident response, establishing consistent practices across teams
  • Define and own SLIs and SLOs for critical services, focusing on reducing mean time to detection (MTTD) and mean time to recovery (MTTR)
  • Lead infrastructure automation efforts using Terraform and enhance CI/CD pipelines with reliability guardrails
Required qualifications:
  • 10+ years of experience in Site Reliability Engineering, DevOps, or infrastructure engineering
  • Demonstrated experience with incident management and production support for high-severity events
  • Expertise in Terraform or comparable infrastructure as code (IaC) tools at scale
  • Hands-on experience with CI/CD pipeline management and cloud platforms (AWS, Azure, or GCP)
  • Proficiency in at least one scripting or programming language (e.g., Python, Go, Bash) for automation

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 47,151 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee