Remote Jobs Sign In

Staff Software Engineer, Reliability

Location: Remote
Compensation: To Be Discussed
Reviewed: Mon, Sep 14, 2026
This job expires in: 30 days

Job Summary

Owning the reliability of the alerting pipeline, the full-time Staff Software Engineer, Reliability will manage end-to-end alert evaluation and delivery, define SLOs and SLIs, and lead incident response efforts, all while working remotely.

Key Responsibilities:
  • Own the reliability of the alerting pipeline from evaluation through delivery, ensuring idempotency and performance under load
  • Define SLOs and SLIs for availability, latency, and delivery, using error budgets to prioritize reliability investments
  • Lead incident response for critical production issues, conducting blameless post-mortems and implementing systemic fixes
Required Qualifications:
  • Experience as a Site Reliability Engineer, DevOps engineer, or in a similar role with expertise in SLOs, SLIs, and incident response
  • Proven track record in designing and operating high-reliability alerting or notification systems at scale
  • Strong experience with Kafka, ensuring webhook reliability and delivery guarantees under load
  • Ability to build observability and automation for high-volume production systems with minimal operational toil
  • Proficiency in Go and/or Python, with experience in container orchestration and multi-cloud environments

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 38,936 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee