Remote Jobs Sign In

Staff Engineer, Inference Optimization

Location: Remote
Compensation: Salary
Reviewed: Thu, Jul 23, 2026
This job expires in: 28 days

Job Summary

To support the AI Inference Optimization team, the full-time Staff Engineer, Inference Optimization will lead architectural decisions to enhance performance, minimize latency, and solve complex bottlenecks in high-performance computing environments, while working remotely.

Key responsibilities
  • Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers
  • Engineer solutions for complex performance issues, including memory and precision management across multi-node GPU clusters
  • Act as a subject matter expert on modern GPU families and their software stacks, advising on hardware procurement and software integration
Required qualifications
  • 5+ years of experience in high-performance computing or AI infrastructure
  • Deep familiarity with the Gen AI landscape and architectural requirements of major model families
  • Hands-on experience with attention-layer optimizations and parallelization strategies across distributed GPU environments
  • Comprehensive understanding of NVIDIA and AMD GPU architectures and their respective software ecosystems
  • Expert-level knowledge of Triton or CUDA, with contributions to relevant open-source projects

COMPLETE JOB DESCRIPTION

The job description is available to subscribers. Subscribe today to get the full benefits of a premium membership with Virtual Vocations. We offer the largest remote database online...