Site Reliability Engineer
Location: Remote
Compensation: Salary
Reviewed: Mon, Aug 17, 2026
This job expires in: 27 days
Job Summary
As a full-time remote Site Reliability Engineer, the successful candidate will ensure the reliability, scalability, and observability of CloudBlue's multi-tenant SaaS platforms, focusing on system stability, performance monitoring, and incident response while collaborating with DevOps and engineering teams.
Key responsibilities
- Define and implement SLIs, SLOs, and error budgets for critical CloudBlue services to ensure reliability and performance
- Design and operate CloudBlue's observability stack using tools such as Datadog and Grafana, while developing actionable alerting strategies
- Lead incident coordination during production incidents, conducting blameless postmortems to drive improvements in system reliability
Required qualifications
- 3+ years of experience as an SRE, DevOps Engineer, or Production Engineer with a strong ownership of production systems
- Proven experience operating highly available, enterprise-grade, multi-tenant SaaS platforms
- Hands-on experience with observability and monitoring tools such as Datadog, Grafana, and Elasticsearch/Kibana
- Solid understanding of Linux, networking, and distributed systems fundamentals
- Experience with containerized environments such as Docker and Kubernetes
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 48,221 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee