Remote Jobs Sign In

Principal ML Performance Engineer

Location: Remote
Compensation: To Be Discussed
Reviewed: Wed, Sep 23, 2026
This job expires in: 30 days

Job Summary

Focused on GPU optimization, the full-time remote Principal ML Performance Engineer will profile and optimize training and inference for advanced machine learning models while managing GPU cluster efficiency and developing benchmarking tools.

Key responsibilities
  • Profile and optimize training and inference for structural and generative models, including transformers and geometric deep learning
  • Manage GPU cluster efficiency on GCP, overseeing scheduling, utilization, and cost reporting
  • Write and tune custom kernels and scale distributed training across multiple nodes
Required qualifications
  • Minimum of 6+ years of experience in ML systems, HPC, or performance engineering, with a BS/MS/PhD in CS, EE, or related field
  • Deep knowledge of PyTorch internals and hands-on experience with profiling and optimizing performance bottlenecks
  • Experience with CUDA and Triton, including proficiency in reading Nsight output
  • Strong proficiency in Python and C++ with experience in distributed training at multi-node scale
  • Ability to demonstrate improvements made to models through optimization efforts

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 40,565 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee