Remote Jobs Sign In

Staff HPC Engineer

Location: Remote
Compensation: To Be Discussed
Reviewed: Thu, Aug 27, 2026
This job expires in: 30 days

Job Summary

Seeking a full-time Staff HPC Engineer, the successful candidate will manage Slurm cluster architecture and multi-tenant scheduling policies while ensuring cluster reliability on both bare-metal and VM-based GPU nodes in a remote capacity.

Key Responsibilities
  • Design, deploy, and operate production Slurm clusters, ensuring high availability and seamless upgrades
  • Model physical fabric for GPU scheduling and enforce multi-tenant scheduling policies, including account management and resource limits
  • Lead the implementation of Slinky on Kubernetes, integrating Slurm with Kubernetes workloads and managing GPU resource allocation
Required Qualifications
  • 8+ years in HPC, systems, or cloud infrastructure engineering, with 4+ years of experience operating production Slurm clusters at scale
  • Deep hands-on expertise with Slurm configuration and management, including authentication and version upgrades on live clusters
  • Strong knowledge of GPU and fabric fundamentals, including NVIDIA drivers and InfiniBand/RoCE technologies
  • Experience with Kubernetes and Slurm-on-Kubernetes stacks, along with automation tools like Terraform and Ansible
  • Proficient in Python and Bash for cluster automation, with additional experience in Go being a plus

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 47,855 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee