Remote Jobs Sign In

Evals Lead, AI Observability

Location: Remote
Compensation: Piece Work
Reviewed: Mon, Oct 05, 2026
This job expires in: 30 days

Job Summary

To support project teams in evaluating LLM-based applications, the remote contract Evals Lead, AI Observability will provide evaluation expertise, develop enablement content, and offer consultancy on evaluation design while ensuring accountability remains with the individual teams.

Key responsibilities
  • Develop enablement and adoption content to guide teams in setting up evaluations, including best practices and methodologies
  • Create reusable evaluation patterns and templates to address common evaluation challenges across projects
  • Consult with project teams on evaluation design, reviewing datasets and measurement strategies while advising on soundness and thresholds
Required qualifications
  • Direct experience designing and running LLM evaluations, including dataset construction and measurement decisions
  • Understanding of LLM-as-judge practices and its calibration against human judgment
  • Knowledge of sampling strategies, annotation guidelines, and managing dataset quality over time
  • Strong grasp of statistical measurement concepts, including sample size and variance
  • Familiarity with LLM application architecture and evaluation tooling such as Langfuse or similar platforms

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 38,620 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee