Senior Site Reliability Engineer
Location: Remote
Compensation: Salary
Reviewed: Mon, Aug 17, 2026
This job expires in: 27 days
Job Summary
Contributing to the deployment and daily operations of large-scale next-generation GPU platforms, the full-time Senior Site Reliability Engineer will manage incidents in GPU clusters, design and implement features in the Base Command Manager product, and validate complex cluster configurations, working both remotely and onsite in Santa Clara.
Key responsibilities
- Contribute to deployments and daily operations of large-scale GPU platforms
- Handle incidents in GPU clusters, bridging operations and development
- Design and implement features in the Base Command Manager product
Required qualifications
- Bachelor's Degree or equivalent experience in Computer Science or related field
- 8+ years of experience in site reliability engineering and/or software development roles
- Fluency in Python
- In-depth knowledge of Linux and networking
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 48,221 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee