Senior HPC Systems Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Fri, Aug 07, 2026
This job expires in: 23 days
Job Summary
To support defense and research programs, the full-time Senior HPC Systems Engineer will build and operate production Slurm clusters, manage hybrid federation of customer-owned clusters, and oversee GPU node operations in a remote environment.
Key responsibilities
- Build and operate production Slurm clusters, including configuration and management of partitions, accounts, and GPU resources
- Connect customer-owned clusters to the control plane, ensuring consistent behavior across different environments
- Manage on-premises hardware provisioning, networking, and fault coordination with site staff or vendors
Required qualifications
- 10+ years of experience operating production Linux systems across multiple distributions
- Proven experience in production Slurm administration, including configuration and debugging
- Experience with parallel or high throughput filesystems and fabric operations such as InfiniBand or RoCE
- Hands-on experience with NVIDIA GPU node operations at multi-node scale
- Familiarity with infrastructure as code frameworks like Ansible or Terraform, and proficiency in scripting languages such as Bash and Python
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 48,944 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee