Senior AI Evaluation Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Sat, Sep 05, 2026
This job expires in: 30 days
Job Summary
Joining a dynamic cross-functional development team, the full-time Senior AI Evaluation Engineer will build evaluation harnesses, create golden datasets, and establish production monitoring for AI systems in a remote work environment.
Key responsibilities
- Build the central evaluation harness and templates for consistent testing across teams
- Create golden datasets with input from business and clinical reviewers, ensuring safety and medical accuracy
- Establish production monitoring for drift and regression, auditing evaluations and reporting metrics monthly
Required qualifications
- 4+ years of experience in ML/LLM evaluation, QA engineering for AI systems, or applied research engineering
- Hands-on experience with evaluation frameworks and statistical rigor on small samples
- Ability to set and defend pass thresholds independently of the development team
- Experience with healthcare or safety-critical evaluation is desirable
- Background in red-teaming is a plus
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 44,401 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee