Remote Jobs Sign In

HPC Infrastructure Engineer

Location: Remote
Compensation: To Be Discussed
Reviewed: Thu, Sep 03, 2026
This job expires in: 30 days

Job Summary

Joining a small research infrastructure team, the full-time remote HPC Infrastructure Engineer will operate and improve NVIDIA GPU clusters, focusing on automation, performance tuning, and hardware management to enhance research throughput and efficiency.

Key responsibilities
  • Operate and enhance the GPU fleet, including provisioning, scheduling, monitoring, and upgrades
  • Build automation to maintain fleet health and streamline operations without human intervention
  • Evaluate and benchmark rented GPU capacity while ensuring security and access control for clusters
Required qualifications
  • Experience managing large-scale Linux server or GPU environments in production
  • Proficient with the NVIDIA stack, including drivers, CUDA, and NCCL
  • Strong skills in automation using Python and/or Bash, with familiarity in IaC tools like Ansible or Terraform
  • Ability to analyze metrics and logs to troubleshoot performance issues effectively
  • Comfortable with hands-on hardware work and datacenter operations

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 44,774 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee