Remote Jobs Sign In

Senior Inference Optimization Engineer

Location: Remote
Compensation: To Be Discussed
Reviewed: Fri, Aug 14, 2026
This job expires in: 29 days

Job Summary

Seeking a full-time Senior Inference Optimization Engineer to work remotely, who will optimize GPU infrastructure, enhance LLM inference performance, and drive cost efficiency for a portfolio company specializing in privacy-first consumer AI.

Key responsibilities:
  • Optimize GPU infrastructure and improve throughput and cost per token for LLM inference workloads
  • Build benchmarking harnesses to identify optimal inference engines and strategies for various workloads
  • Evaluate and implement emerging inference optimization techniques and hardware solutions
Required qualifications:
  • 5+ years of experience in performance optimization or HPC with knowledge of GPU architecture and parallel programming
  • Hands-on experience with production LLM inference engines operating at high volume
  • Demonstrated experience in LLM inference optimization techniques, including continuous batching and KV cache management
  • Fluency in GPU profiling tools such as Nsight Systems and PyTorch Profiler
  • Proficiency in programming languages such as Python, Rust, or Go; C++/CUDA is a strong plus

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 46,074 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee