Senior Site Reliability Engineer
Location: Remote
Compensation: Salary
Reviewed: Thu, Sep 10, 2026
This job expires in: 30 days
Job Summary
Owning the reliability of production SaaS services in a remote full-time capacity, the Senior Site Reliability Engineer will manage availability, performance, and capacity, while automating responses and leading incident resolution in a FedRAMP High environment.
Key responsibilities:
- Define service-level indicators and objectives, prioritizing reliability work with engineering and product teams
- Build and tune monitoring systems in Datadog and Azure Monitor to enhance detection and reduce alert noise
- Participate in on-call rotations, leading incident responses and writing post-incident reviews to drive preventive actions
Required qualifications:
- 8+ years in Site Reliability Engineering, DevOps, or related fields for a SaaS product
- Hands-on experience with Azure services, including AKS, App Service, and Azure SQL
- Proficiency in using an observability platform like Datadog for metrics and logs
- Experience with Kubernetes in production environments
- Familiarity with infrastructure as code using Terraform and CI/CD pipelines
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 42,189 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee