Remote Jobs Sign In

Senior Site Reliability Engineer

Location: Remote
Compensation: To Be Discussed
Reviewed: Mon, Aug 31, 2026
This job expires in: 30 days

Job Summary

Owning production readiness and daily operations for a self-hosted Langfuse platform on AWS EKS, the contract Senior Site Reliability Engineer will focus on reliability, observability, incident management, automation, and scaling across various technologies including Kubernetes, ClickHouse, PostgreSQL, and Terraform in a remote work environment.

Key responsibilities
  • Manage production operations and reliability for the Langfuse platform on AWS EKS
  • Define and implement SLIs, SLOs, dashboards, alerts, and operational runbooks
  • Support incident triage, troubleshooting, rollback, escalation, and post-incident reviews
Required qualifications
  • Strong hands-on experience managing production Kubernetes environments, preferably AWS EKS
  • Experience with AWS services, including IAM/IRSA, S3, and networking
  • Hands-on experience operating ClickHouse, including replication and performance troubleshooting
  • Experience with Terraform and Infrastructure as Code
  • Strong experience with Argo CD, GitOps, and CI/CD pipelines

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 43,976 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee