Staff Software Engineer, Reliability
Location: Remote
Compensation: To Be Discussed
Reviewed: Mon, Sep 14, 2026
This job expires in: 30 days
Job Summary
Owning the reliability of the alerting pipeline, the full-time Staff Software Engineer, Reliability will manage end-to-end alert evaluation and delivery, define SLOs and SLIs, and lead incident response efforts, all while working remotely.
Key Responsibilities:
- Own the reliability of the alerting pipeline from evaluation through delivery, ensuring idempotency and performance under load
- Define SLOs and SLIs for availability, latency, and delivery, using error budgets to prioritize reliability investments
- Lead incident response for critical production issues, conducting blameless post-mortems and implementing systemic fixes
Required Qualifications:
- Experience as a Site Reliability Engineer, DevOps engineer, or in a similar role with expertise in SLOs, SLIs, and incident response
- Proven track record in designing and operating high-reliability alerting or notification systems at scale
- Strong experience with Kafka, ensuring webhook reliability and delivery guarantees under load
- Ability to build observability and automation for high-volume production systems with minimal operational toil
- Proficiency in Go and/or Python, with experience in container orchestration and multi-cloud environments
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 38,936 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee