Staff HPC Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Thu, Aug 27, 2026
This job expires in: 30 days
Job Summary
Seeking a full-time Staff HPC Engineer, the successful candidate will manage Slurm cluster architecture and multi-tenant scheduling policies while ensuring cluster reliability on both bare-metal and VM-based GPU nodes in a remote capacity.
Key Responsibilities
- Design, deploy, and operate production Slurm clusters, ensuring high availability and seamless upgrades
- Model physical fabric for GPU scheduling and enforce multi-tenant scheduling policies, including account management and resource limits
- Lead the implementation of Slinky on Kubernetes, integrating Slurm with Kubernetes workloads and managing GPU resource allocation
Required Qualifications
- 8+ years in HPC, systems, or cloud infrastructure engineering, with 4+ years of experience operating production Slurm clusters at scale
- Deep hands-on expertise with Slurm configuration and management, including authentication and version upgrades on live clusters
- Strong knowledge of GPU and fabric fundamentals, including NVIDIA drivers and InfiniBand/RoCE technologies
- Experience with Kubernetes and Slurm-on-Kubernetes stacks, along with automation tools like Terraform and Ansible
- Proficient in Python and Bash for cluster automation, with additional experience in Go being a plus
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 47,855 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee