Senior Site Reliability Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Wed, Sep 02, 2026
This job expires in: 28 days
Job Summary
Owning the reliability, performance, and resilience of cloud infrastructure, the remote Senior Site Reliability Engineer will lead incident response, define SLOs, and automate operational processes to enhance the efficiency of Garner's AI/ML workloads.
Key responsibilities
- Own the end-to-end reliability and performance of cloud environments, defining and measuring SLOs across critical services
- Lead incident response efforts and conduct root cause analyses to ensure resolution and infrastructure change review
- Build and maintain monitoring and observability systems to proactively detect and resolve issues
Required qualifications
- 4+ years of experience in operating production cloud infrastructure in an SRE, DevOps, or platform engineering role
- Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred)
- Strong track record in production observability, including defining SLOs and leading incident responses
- Proficiency in Python or Go for infrastructure automation
- Experience with security and compliance in regulated environments such as HIPAA is a plus
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 44,401 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee