Evaluation Lead for AI
Location: Remote
Compensation: Piece Work
Reviewed: Fri, Sep 25, 2026
This job expires in: 30 days
Job Summary
To support project teams in effectively evaluating LLM applications and agents, the remote contract Evaluation Lead for AI will provide evaluation expertise, develop reusable evaluation patterns, and offer consultancy on evaluation design while ensuring accountability remains with the respective teams.
Key responsibilities
- Develop enablement and adoption content for setting up evaluations, including written guidance and best practices
- Create reusable evaluation patterns and templates addressing various evaluation challenges across projects
- Consult with project teams on evaluation design, reviewing datasets and recommending measurement thresholds
Required qualifications
- Direct experience in designing and running LLM evaluations, including dataset construction and measurement decisions
- Knowledge of LLM application architecture and its limitations, including retrieval and prompt management
- Strong understanding of measurement concepts such as sample size, statistical significance, and variance
- Experience with LLM observability and evaluation tools like Langfuse or similar platforms is a plus
- Proficient written and verbal communication skills
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 41,934 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee