×
Register Here to Apply for Jobs or Post Jobs. X

Research Crawling Engineer

Job in Moreno Valley, Riverside County, California, 92551, USA
Listing for: Wintermeyer Ventures
Full Time position
Listed on 2026-09-30
Job specializations:
  • IT/Tech
    Data Engineering
Salary/Wage Range or Industry Benchmark: 160000 - 250000 USD Yearly USD 160000.00 250000.00 YEAR
Job Description & How to Apply Below

About the Role

This is a hands-on engineering role focused on building and operating large-scale web crawlers and data acquisition systems that power dataset creation for frontier AI model training. You will work as an extension of AI research teams, helping to collect, clean, and curate web data at a scale that few organizations can match. The work directly shapes the pretraining and inference pipelines used by some of the most advanced AI labs in the world.

What

You'll Do
  • Build and maintain large-scale web crawlers across diverse domains including social media, travel, and multi-language sites.
  • Design high-throughput, fault-tolerant data collection systems capable of handling millions to billions of URLs per day.
  • Navigate anti-bot systems, rate limits, and JavaScript-heavy sites, finding creative solutions when standard protocols fall short.
  • Develop pipelines for cleaning, deduplication, filtering, and normalization of web data at TB to PB scale.
  • Construct and maintain datasets for research and model training in close collaboration with research teams.
  • Monitor crawl performance, coverage, and data quality, iterating quickly as web environments change.
  • Optimize infrastructure for cost, latency, and reliability across cloud and bare-metal environments.
What We're Looking For
  • 3 or more years building and operating web crawlers at scale (2 or more years considered for candidates with a PhD).
  • Proficiency in one or more of:
    Go, Rust, Python, Java, or C++.
  • Experience running data pipelines at TB or greater scale.
  • Deep knowledge of HTTP, networking, and browser behavior.
  • Hands‑on experience with distributed systems or parallel processing.
  • Experience with headless browsers such as Playwright, Puppeteer, or Chrome Dev Tools Protocol.
  • Familiarity with proxy systems, IP rotation, or request orchestration.
  • Experience with data quality evaluation, scoring, or benchmarking at scale.
  • Experience running crawling or data workloads on cloud platforms (AWS, GCP) or bare-metal infrastructure.
  • Background in NLP pipelines, ML dataset curation, or AI lab work is a strong plus.
Compensation & Benefits

Salary range: $160,000 to $250,000 USD annually. Visa sponsorship is not available for this role.

Location

This role is fully remote. The primary location is Los Angeles, CA, United States, though candidates based in other major US cities are welcome.

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary