Site Reliability Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Thu, Sep 24, 2026
This job expires in: 30 days
Job Summary
Focusing on the reliability and scalability of CloudBlue's multi-tenant SaaS platforms, the full-time Site Reliability Engineer will work remotely to improve system stability and performance through monitoring, incident response, and collaboration with DevOps and engineering teams.
Key responsibilities
- Define and implement SLIs, SLOs, and error budgets for critical services to ensure reliability and performance
- Design and operate observability stacks using tools such as Datadog and Grafana, while developing actionable alerting strategies
- Lead incident coordination during production incidents and drive improvements to reduce incident frequency and customer impact
Required qualifications
- 3+ years of experience as an SRE, DevOps Engineer, or Production Engineer
- Proven experience operating highly available, enterprise-grade, multi-tenant SaaS platforms
- Hands-on experience with observability and monitoring tools like Datadog and Grafana
- Solid understanding of Linux, networking, and distributed systems fundamentals
- Experience with containerized environments such as Docker and Kubernetes
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 41,975 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee