Staff Engineer, Inference Optimization
Location: Remote
Compensation: Salary
Reviewed: Thu, Jul 23, 2026
This job expires in: 28 days
Job Summary
To support the AI Inference Optimization team, the full-time Staff Engineer, Inference Optimization will lead architectural decisions to enhance performance, minimize latency, and solve complex bottlenecks in high-performance computing environments, while working remotely.
Key responsibilities
- Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers
- Engineer solutions for complex performance issues, including memory and precision management across multi-node GPU clusters
- Act as a subject matter expert on modern GPU families and their software stacks, advising on hardware procurement and software integration
Required qualifications
- 5+ years of experience in high-performance computing or AI infrastructure
- Deep familiarity with the Gen AI landscape and architectural requirements of major model families
- Hands-on experience with attention-layer optimizations and parallelization strategies across distributed GPU environments
- Comprehensive understanding of NVIDIA and AMD GPU architectures and their respective software ecosystems
- Expert-level knowledge of Triton or CUDA, with contributions to relevant open-source projects
COMPLETE JOB DESCRIPTION
The job description is available to subscribers. Subscribe today to get the full benefits of a premium membership with Virtual Vocations. We offer the largest remote database online...