Senior Site Reliability Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Fri, Sep 11, 2026
This job expires in: 30 days
Job Summary
Driving the reliability and observability strategy for a financial data platform, the full-time Senior Site Reliability Engineer will manage the reliability roadmap, enhance foundational systems, and implement scalable incident processes while working remotely in the United States.
Key responsibilities
- Own the reliability roadmap, setting SLOs and prioritizing work to enhance platform observability and performance
- Harden the financial data platform to ensure enterprise-grade availability and recoverability, improving deployment paths and processing efficiency
- Build and manage an incident and review process that maintains high reliability as the team and systems grow
Required qualifications
- Proven experience designing and operating reliable production systems at scale
- Deep expertise with observability tools such as Datadog, Prometheus/Grafana, and OpenTelemetry
- Strong background in defining SLIs, SLOs, and error budgets, translating them into business-level KPIs
- Hands-on experience with data platforms like Snowflake and enterprise integration layers such as MuleSoft
- Experience in incident management using tools like Incident.io and ServiceNow, with a track record of effective on-call practices
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 42,189 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee