Senior Site Reliability Engineer
Location: Remote
Compensation: Salary
Reviewed: Tue, Sep 22, 2026
This job expires in: 30 days
Job Summary
To support a growing team focused on AI and automation, the full-time Senior Site Reliability Engineer will design and maintain observability systems, operate Kubernetes-based platforms, and build AI agents to enhance operational efficiency, all while working remotely.
Key responsibilities
- Participate in on-call rotations to diagnose and resolve production issues using established runbooks and playbooks
- Design, build, and maintain observability dashboards and alerting systems based on Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
- Collaborate with product engineering teams to review architecture and infrastructure decisions, ensuring reliability and scalability
Required qualifications
- Strong, hands-on experience with Kubernetes as a system
- Practical experience using AI tools and agents for automating infrastructure and reliability tasks
- Solid grounding in AWS or Azure, including networking fundamentals
- Deep experience with modern observability stacks such as OpenTelemetry, Prometheus, or Grafana
- 8-10+ years of relevant hands-on experience in Site Reliability Engineering
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 39,143 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee