Remote Jobs Sign In

Research Engineer - Web Crawlers

Location: Remote
Compensation: To Be Discussed
Reviewed: Fri, Aug 28, 2026
This job expires in: 28 days

Job Summary

Focused on large-scale web crawling for frontier AI models, the full-time remote Research Engineer - Web Crawlers will build and operate distributed web crawlers, solve complex data extraction challenges, and design targeted crawling pipelines to create high-quality training datasets.

Key responsibilities
  • Build and operate large-scale, distributed web crawlers to efficiently discover and extract data from billions of web pages
  • Solve complex challenges related to content extraction, deduplication, and data freshness in web crawling
  • Design targeted crawling pipelines to identify and convert high-value data sources into clean, training-ready datasets
Required qualifications
  • Hands-on experience in building and scaling web crawlers or scraping systems, particularly for machine learning training data
  • Strong engineering skills in distributed systems at scale, such as Kubernetes or queue-based architectures
  • Ability to autonomously evaluate the quality, coverage, and compliance of crawled data
  • Experience in creating tooling to measure and monitor crawled data effectively

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 45,457 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee