Research Engineer - Web Crawlers
Location: Remote
Compensation: To Be Discussed
Reviewed: Fri, Aug 28, 2026
This job expires in: 28 days
Job Summary
Focused on large-scale web crawling for frontier AI models, the full-time remote Research Engineer - Web Crawlers will build and operate distributed web crawlers, solve complex data extraction challenges, and design targeted crawling pipelines to create high-quality training datasets.
Key responsibilities
- Build and operate large-scale, distributed web crawlers to efficiently discover and extract data from billions of web pages
- Solve complex challenges related to content extraction, deduplication, and data freshness in web crawling
- Design targeted crawling pipelines to identify and convert high-value data sources into clean, training-ready datasets
Required qualifications
- Hands-on experience in building and scaling web crawlers or scraping systems, particularly for machine learning training data
- Strong engineering skills in distributed systems at scale, such as Kubernetes or queue-based architectures
- Ability to autonomously evaluate the quality, coverage, and compliance of crawled data
- Experience in creating tooling to measure and monitor crawled data effectively
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 45,457 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee