Site Reliability Engineering Lead
Location: Remote
Compensation: Salary
Reviewed: Thu, Sep 17, 2026
This job expires in: 30 days
Job Summary
Leading the Site Reliability Engineering team, the full-time Principal SRE will own the reliability and operational health of a cloud-native automotive AI platform while mentoring team members, defining reliability roadmaps, and serving as a Tier 2 technical escalation point for major incidents in a remote work environment.
Key responsibilities
- Set technical direction and priorities for the team while mentoring and developing members across multiple locations
- Own and drive the team's reliability roadmap, defining SLI/SLO/SLA frameworks to ensure contracted availability targets
- Serve as the Tier 2 technical escalation point for major incidents and lead continuous improvement initiatives for production readiness
Required qualifications
- 8+ years of hands-on experience in site reliability, DevOps, or cloud platform roles, including leadership experience
- Proficiency with container orchestration frameworks such as Kubernetes and Docker
- Experience with public cloud platforms, primarily Azure, as well as AWS and Google Cloud
- Familiarity with observability tools and CI/CD pipelines, including infrastructure-as-code practices
- Strong UNIX/Linux background, including system configuration and performance debugging
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 41,606 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee