Remote Jobs Sign In

Senior Site Reliability Engineer

Location: Remote
Compensation: Salary
Reviewed: Tue, Aug 18, 2026
This job expires in: 29 days

Job Summary

As a full-time Senior Cluster Site Reliability Engineer working remotely, the successful candidate will manage the scaling of research compute clusters, ensure high uptime and reliability, and support both on-prem and cloud infrastructure to provide a robust HPC platform for researchers.

Key responsibilities
  • Be a first responder in the event of cluster outages or issues, triaging and resolving urgent problems as they arise
  • Ensure a high degree of cluster uptime and track SLAs to quantify reliability
  • Develop robust metrics and observability for cluster health, informing operational improvements and policy design
Required qualifications
  • 5+ years of experience in SRE or DevOps roles, preferably as a senior engineer or tech lead
  • Knowledge of HPC/batch compute frameworks and machine learning training systems
  • Ability to develop scripts in a common scripting language such as Python or Ruby
  • Familiarity with infrastructure-as-code and configuration management tools like Terraform or Ansible
  • Experience with cloud infrastructure, specifically AWS or GCP

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 47,838 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee