Senior Software Engineer, GPU Infrastructure
Location: Remote
Compensation: Salary
Reviewed: Fri, Sep 11, 2026
This job expires in: 30 days
Job Summary
Operating the Kubernetes GPU fleet and managing multi-tenancy, the full-time Senior Software Engineer, GPU Infrastructure will focus on maintaining large-scale experiments, ensuring fault tolerance, and enhancing platform security in a remote work environment.
Key responsibilities
- Handle day-to-day operations of the Kubernetes GPU fleet, including capacity planning and node lifecycle management
- Design and implement batch scheduling and storage solutions for the GPU cluster, ensuring efficient resource allocation
- Collaborate directly with research teams to address infrastructure challenges and improve platform functionality
Required qualifications
- 3+ years of experience in systems or infrastructure engineering on production Linux, specifically with GPU or large-scale batch platforms
- Proven experience managing production Kubernetes for GPU workloads, including batch layers and resource management
- Familiarity with infrastructure as code tools such as Terraform or Ansible, and monitoring solutions like Prometheus
- Strong programming skills in languages commonly used for infrastructure, such as Python, Go, Rust, or C++
- Ability to write clear documentation for engineers, researchers, and providers, including design documents and incident reports
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 42,237 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee