Technical Lead - GPU Infrastructure
Location: Remote
Compensation: To Be Discussed
Reviewed: Fri, Sep 11, 2026
This job expires in: 30 days
Job Summary
Leading a distributed engineering team, the full-time Technical Lead - GPU Infrastructure will own the architecture and delivery of a GPU compute platform, focusing on building and operating a managed Slurm service and Kubernetes control plane, all while working remotely.
Key responsibilities
- Own the end-to-end platform architecture, including design proposals and current baseline maintenance
- Lead and manage a distributed team across various engineering disciplines, ensuring adherence to engineering standards and performance growth
- Design, build, and operate a managed Slurm scheduling layer for research users, ensuring efficient resource allocation and system health
Required qualifications
- Eight or more years of hands-on engineering experience, with at least three years in a leadership role building infrastructure platforms
- Proven experience with Slurm at scale, including operational management of HPC or GPU training clusters
- Deep knowledge of GPU fleet operation on bare metal, including NVIDIA driver and CUDA lifecycle management
- Strong background in production Kubernetes operations, including control plane management and multi-tenancy design
- Working fluency in JavaScript and Node.js for architecture decision-making and code review
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 42,189 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee