Remote Jobs Sign In

Senior Site Reliability Engineer

Location: Remote
Compensation: Salary
Reviewed: Tue, Sep 15, 2026
This job expires in: 30 days

Job Summary

To support the DGX Cloud team, the full-time Senior Site Reliability Engineer will manage and optimize large-scale Kubernetes clusters, ensuring high performance and reliability for AI workloads in a remote environment.

Key responsibilities
  • Build, implement, and support operational aspects of Kubernetes clusters with a focus on performance, monitoring, and alerting
  • Define SLOs/SLIs, monitor system health, and streamline reporting processes
  • Lead incident response efforts, including triage and root-cause analysis of high-severity incidents
Required qualifications
  • BS in Computer Science or related technical field, or equivalent experience
  • 8+ years of experience operating production services
  • Expert-level knowledge of Kubernetes administration and microservices architecture
  • Experience with infrastructure automation tools such as Terraform or Ansible
  • Proficiency in at least one high-level programming language, such as Python or Go

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 39,019 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee