Site Reliability Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Fri, Oct 09, 2026
This job expires in: 30 days
Job Summary
Owning the operational health of provider supply, the full-time Site Reliability Engineer will build monitoring systems, improve detection and failover processes, and manage incident responses while working remotely.
Key responsibilities
- Build and maintain monitoring for provider health, including latency, throughput, and error rates, while setting SLOs and alerts
- Enhance detection of degraded endpoints and collaborate with the routing team for automatic traffic shifting
- Lead incident response for provider issues, including triage, communication, and postmortem analysis
Required qualifications
- 4+ years of experience in SRE, production engineering, or infrastructure roles with high-traffic systems
- Proficiency in observability tools and practices, including metrics, logs, and alerting
- Strong software engineering skills, preferably in TypeScript and/or Python
- Experience with distributed systems and their failure modes
- Effective communication skills for incident management and external partner interactions
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 42,491 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee