Remote Jobs Sign In

Senior Site Reliability Engineer

Location: Remote
Compensation: Salary
Reviewed: Mon, Aug 17, 2026
This job expires in: 27 days

Job Summary

Contributing to the deployment and daily operations of large-scale next-generation GPU platforms, the full-time Senior Site Reliability Engineer will manage incidents in GPU clusters, design and implement features in the Base Command Manager product, and validate complex cluster configurations, working both remotely and onsite in Santa Clara.

Key responsibilities
  • Contribute to deployments and daily operations of large-scale GPU platforms
  • Handle incidents in GPU clusters, bridging operations and development
  • Design and implement features in the Base Command Manager product
Required qualifications
  • Bachelor's Degree or equivalent experience in Computer Science or related field
  • 8+ years of experience in site reliability engineering and/or software development roles
  • Fluency in Python
  • In-depth knowledge of Linux and networking

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 48,221 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee