Junior HPC Systems Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Fri, Aug 07, 2026
This job expires in: 23 days
Job Summary
Seeking a full-time Junior HPC Systems Engineer to work remotely, who will monitor cluster health, manage user accounts and allocations, and perform node lifecycle management while gaining experience in supercomputing operations on production systems.
Key responsibilities
- Monitor and respond to cluster health, node state, and alerting for failures and stuck jobs
- Manage user accounts, groups, and Slurm allocations across various computing environments
- Conduct health checks on GPU and CPU nodes, and escalate faults as necessary
Required qualifications
- 2+ years of hands-on Linux administration experience (RHEL, Rocky, Alma, Debian, Ubuntu)
- Understanding of HPC fundamentals, including batch scheduling and shared filesystems
- Proficiency in Bash and Python scripting for operational tasks
- Practical knowledge of networking concepts such as DNS, routing, and firewalls
- U.S. citizenship and eligibility for a Secret clearance
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 48,944 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee