Kubernetes Site Reliability Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Thu, Aug 27, 2026
This job expires in: 27 days
Job Summary
Managing the control plane for AI-operated GPU cloud infrastructure, the full-time Kubernetes Site Reliability Engineer will design, deploy, and operate production Kubernetes clusters optimized for GPU workloads, ensuring automated remediation and multi-tenant isolation in a remote role based in San Jose, CA or Austin, TX.
Key responsibilities
- Design and maintain production Kubernetes clusters optimized for GPU workloads at scale
- Implement topology-aware scheduling and manage GPU-specific resource allocation
- Automate incident management processes, including runbook automation and monitoring stack integration
Required qualifications
- 5+ years of experience in Kubernetes operations, with at least 2 years managing GPU workloads
- Deep understanding of Nvidia GPU operator and GPU scheduling in Kubernetes
- Experience building multi-tenant Kubernetes platforms with strong isolation guarantees
- Proficiency in Terraform, Helm, and GitOps workflows
- Strong programming skills in Go or Python for operator/CRD development
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 45,457 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee