Remote Jobs Sign In

Kubernetes Site Reliability Engineer

Location: Remote
Compensation: To Be Discussed
Reviewed: Thu, Aug 27, 2026
This job expires in: 27 days

Job Summary

Managing the control plane for AI-operated GPU cloud infrastructure, the full-time Kubernetes Site Reliability Engineer will design, deploy, and operate production Kubernetes clusters optimized for GPU workloads, ensuring automated remediation and multi-tenant isolation in a remote role based in San Jose, CA or Austin, TX.

Key responsibilities
  • Design and maintain production Kubernetes clusters optimized for GPU workloads at scale
  • Implement topology-aware scheduling and manage GPU-specific resource allocation
  • Automate incident management processes, including runbook automation and monitoring stack integration
Required qualifications
  • 5+ years of experience in Kubernetes operations, with at least 2 years managing GPU workloads
  • Deep understanding of Nvidia GPU operator and GPU scheduling in Kubernetes
  • Experience building multi-tenant Kubernetes platforms with strong isolation guarantees
  • Proficiency in Terraform, Helm, and GitOps workflows
  • Strong programming skills in Go or Python for operator/CRD development

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 45,457 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee