Senior Site Reliability Engineer
Location: Remote
Compensation: Salary
Reviewed: Tue, Aug 18, 2026
This job expires in: 29 days
Job Summary
As a full-time Senior Cluster Site Reliability Engineer working remotely, the successful candidate will manage the scaling of research compute clusters, ensure high uptime and reliability, and support both on-prem and cloud infrastructure to provide a robust HPC platform for researchers.
Key responsibilities
- Be a first responder in the event of cluster outages or issues, triaging and resolving urgent problems as they arise
- Ensure a high degree of cluster uptime and track SLAs to quantify reliability
- Develop robust metrics and observability for cluster health, informing operational improvements and policy design
Required qualifications
- 5+ years of experience in SRE or DevOps roles, preferably as a senior engineer or tech lead
- Knowledge of HPC/batch compute frameworks and machine learning training systems
- Ability to develop scripts in a common scripting language such as Python or Ruby
- Familiarity with infrastructure-as-code and configuration management tools like Terraform or Ansible
- Experience with cloud infrastructure, specifically AWS or GCP
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 47,838 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee