Principal Site Reliability Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Wed, Sep 16, 2026
This job expires in: 30 days
Job Summary
Setting the reliability strategy for the platform, the full-time Principal Site Reliability Engineer will define deployment and operational standards for distributed systems, ensuring reliability and automation across customer environments while working remotely.
Key responsibilities
- Own the reliability architecture of the platform, including deployment topology and automation for reproducible environments
- Define service level objectives and drive initiatives to meet them in collaboration with product and engineering leadership
- Lead major incidents and postmortems, ensuring corrective actions are implemented effectively
Required qualifications
- 10+ years in infrastructure, SRE, or platform engineering with experience in large-scale distributed systems
- Expertise in designing and troubleshooting distributed systems and making reliability decisions under growth
- Deep experience with at least one major public cloud provider and proficiency in container orchestration
- Strong scripting and automation skills, along with familiarity in reading and debugging application code
- Experience operating within enterprise security and compliance frameworks
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 40,513 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee