Senior Inference Engineer
Location: Remote
Compensation: Salary
Reviewed: Tue, Aug 11, 2026
This job expires in: 22 days
Job Summary
Partnering directly with the CTO, the full-time Senior Inference Engineer will build and optimize the platform's inference layer, deploying models to production and managing the entire query-to-response pipeline in a remote environment.
Key responsibilities
- Deploying models to production and owning the full pipeline from query to response
- Building and operating serving infrastructure using frameworks like vLLM, SGLang, or TensorRT-LLM
- Leading the engineering direction for inference, collaborating with Product to shape future developments
Required qualifications
- Experience deploying and serving LLMs using frameworks such as vLLM, SGLang, or TensorRT-LLM
- Hands-on knowledge of inference optimization techniques like quantization, batching, caching, and routing
- Strong programming skills in Python or Golang with real production code experience
- Ability to communicate technical concepts clearly to both technical and non-technical stakeholders
- Familiarity with containerized environments (Docker, Kubernetes) is a plus
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 47,838 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee