Principal ML Performance Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Wed, Sep 23, 2026
This job expires in: 30 days
Job Summary
Focused on GPU optimization, the full-time remote Principal ML Performance Engineer will profile and optimize training and inference for advanced machine learning models while managing GPU cluster efficiency and developing benchmarking tools.
Key responsibilities
- Profile and optimize training and inference for structural and generative models, including transformers and geometric deep learning
- Manage GPU cluster efficiency on GCP, overseeing scheduling, utilization, and cost reporting
- Write and tune custom kernels and scale distributed training across multiple nodes
Required qualifications
- Minimum of 6+ years of experience in ML systems, HPC, or performance engineering, with a BS/MS/PhD in CS, EE, or related field
- Deep knowledge of PyTorch internals and hands-on experience with profiling and optimizing performance bottlenecks
- Experience with CUDA and Triton, including proficiency in reading Nsight output
- Strong proficiency in Python and C++ with experience in distributed training at multi-node scale
- Ability to demonstrate improvements made to models through optimization efforts
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 40,565 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee