Inference Performance Engineer
Location: Remote
Compensation: Salary
Reviewed: Thu, Aug 06, 2026
This job expires in: 18 days
Job Summary
Optimizing AI inference configurations, the full-time Inference Performance Engineer will enhance performance on large-scale benchmarks by developing reusable workflows and methodologies, profiling workloads, and collaborating with cross-functional teams, with opportunities for both onsite and remote work.
Key responsibilities
- Develop reusable skills and workflows for autonomous AI agents to optimize performance in AI inference workloads
- Measure and enhance throughput-per-GPU and user interactivity through various optimization techniques and configurations
- Collaborate with multiple teams to translate profiling insights into effective performance improvements across NVIDIA's GPU platforms
Required qualifications
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Math, or a related field, or equivalent experience
- 3+ years of relevant engineering experience in AI model execution optimization
- Extensive knowledge of performance benchmarking and profiling GPU workloads using tools like Nsight Systems and PyTorch profiler
- Strong Python engineering skills and experience with large C++/CUDA codebases
- Rigorous experimental methodology for controlled comparisons and reproducible benchmarks
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 44,899 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee