GPU Cluster Engineer
Location: Remote
Compensation: Salary
Reviewed: Thu, Oct 08, 2026
This job expires in: 30 days
Job Summary
To build high-performance GPU clusters, the contract GPU Cluster Infrastructure Engineer will manage design reviews, lead acceptance testing, and establish operational foundations for GPU capacity, working remotely or onsite in New York or San Francisco.
Key responsibilities
- Review cluster designs and bills of materials to identify gaps before hardware ordering
- Lead acceptance testing, including validating cabling, running burn-in, and ensuring vendor deliverables
- Integrate hardware, fabric, and storage telemetry into the observability stack with alerting and automated health checks
Required qualifications
- 3+ years of experience building and operating NVIDIA HGX or DGX clusters in production settings
- Hands-on experience with InfiniBand, including subnet management and fabric bring-up
- Familiarity with GPU node bring-up processes and parallel storage solutions like WEKA or Lustre
- Ability to work effectively both on-site and remotely, including directing remote hands
- Experience with automation tools such as Ansible and monitoring tools like Prometheus/Grafana is a plus
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 41,328 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee