Staff ML Platform Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Fri, Sep 25, 2026
This job expires in: 30 days
Job Summary
To support the Data & ML Platform organization, the full-time Staff ML Platform Engineer will set technical direction for ML training and serving, own architectural solutions for LLM-endpoint serving, and evolve the paved-road framework to enable Data Science teams to deploy models effectively in a remote environment.
Key responsibilities
- Set technical direction across ML training, serving, and observability while addressing complex infrastructure problems
- Own and evolve the paved-road framework to facilitate seamless production workflows for Data Science teams
- Lead architecture for LLM-endpoint serving, ensuring efficient management of latency, cost, and PHI-safe routing
Required qualifications
- 10+ years of software engineering experience, with at least 3 years in enterprise-scale ML platforms
- Hands-on experience with Databricks, Amazon SageMaker, and MLflow or similar systems
- Proficiency in Java (or JVM equivalent) and Python, with expertise in Apache Spark
- Strong knowledge of AWS services relevant to ML workloads, including networking and GPU compute
- Experience with Terraform, containers, Kubernetes, and CI/CD for ML applications
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 42,033 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee