Remote Jobs Sign In

Inference Infrastructure Architect

Location: Remote
Compensation: To Be Discussed
Reviewed: Mon, Sep 21, 2026
This job expires in: 30 days

Job Summary

Seeking a remote Inference Infrastructure Architect, the full-time position will focus on optimizing a global fleet for AI-agent traffic, ensuring efficient throughput and latency while managing serverless inference deployments for both internal and external clients.

Key responsibilities
  • Operate and expand the GPU fleet to maximize inference throughput and minimize costs
  • Develop serverless serving pools and dedicated model deployments for enterprise clients
  • Implement observability and capacity planning strategies to enhance performance and reliability
Required qualifications
  • Proven experience with production LLM serving under traffic and latency constraints
  • Hands-on expertise with Kubernetes on GPU fleets, including GPU Operator and topology-aware scheduling
  • Deep operational knowledge of vLLM or SGLang, including deployment and tuning
  • Strong background in performance engineering and system-level metrics analysis
  • Proficiency in Python and Go for automation, with a solid understanding of Linux and networking

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 38,952 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee