Site Reliability Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Tue, Sep 15, 2026
This job expires in: 30 days
Job Summary
The full-time remote Site Reliability Engineer will manage observability and incident response for a complex distributed system, influencing technical direction and mentoring engineers while ensuring system reliability and performance.
Key responsibilities
- Own the technical direction of the observability stack, defining instrumentation standards for Java and Node.js services
- Establish SLIs, SLOs, and error budgets, partnering with engineering and product teams to drive engineering decisions
- Lead major incident response as a senior incident commander and conduct blameless postmortems with actionable follow-through
Required qualifications
- 5+ years in SRE, infrastructure, or platform engineering with experience in large-scale production systems
- Deep production experience with Kubernetes, preferably GKE, and proficiency in debugging under pressure
- Strong observability background with OpenTelemetry, Prometheus, and centralized logging tools
- Hands-on experience with stateful services in production, including PostgreSQL, MongoDB Atlas, and RabbitMQ
- Proven track record in leading incident response and SLO programs that positively impacted engineering behavior
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 39,019 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee