Director of Site Reliability Engineering
Location: Remote
Compensation: To Be Discussed
Reviewed: Sat, Sep 26, 2026
This job expires in: 30 days
Job Summary
Building and leading the Site Reliability Engineering function from the ground up, the full-time Director of Site Reliability Engineering will own critical infrastructure across colocation, on-premises, and cloud environments while managing a team and driving reliability and cost efficiency in a hybrid work setting.
Key responsibilities
- Establish the SRE function, defining its charter and hiring a team of SRE engineers
- Own 24/7 reliability across various platforms, designing for failure domains and implementing strict change control
- Drive Infrastructure as Code automation and manage the full observability stack for SLO visibility
Required qualifications
- Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field; 15+ years in SRE or infrastructure engineering
- 5+ years of experience leading SRE or infrastructure engineering teams, with a proven track record of establishing SRE as a discipline
- Deep expertise in Linux systems and experience operating colocation and on-premises hardware at scale
- Hands-on fluency with Infrastructure as Code tools such as Terraform and Ansible at production scale
- Strong scripting skills in Python and/or Go for automation and production services
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 41,659 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee