Staff Engineer, AI Inference
Location: Remote
Compensation: Salary
Reviewed: Thu, Jul 23, 2026
This job expires in: 24 days
Job Summary
To ensure industry-leading performance for inference services, the full-time Staff Engineer, AI Inference will lead architectural decisions, optimize performance at the GPU kernel level, and implement cutting-edge techniques in a remote work environment.
Key responsibilities
- Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers
- Engineer solutions for complex performance issues, including attention layer optimizations and advanced parallelization across multi-node GPU clusters
- Act as a subject matter expert on modern GPU families and their software stacks, advising on hardware procurement and software integration
Required qualifications
- 5+ years of experience in high-performance computing or AI infrastructure
- Deep familiarity with the Gen AI landscape, including major model families
- Hands-on experience with attention-layer optimizations and parallelization strategies in distributed GPU environments
- Comprehensive understanding of NVIDIA and AMD GPU architectures and their software ecosystems
- Expert-level proficiency in Triton or CUDA, with experience in contributing to open-source software projects
COMPLETE JOB DESCRIPTION
The job description is available to subscribers. Subscribe today to get the full benefits of a premium membership with Virtual Vocations. We offer the largest remote database online...