More jobs:
Research Crawling Engineer
Job in
Modesto, Stanislaus County, California, 95350, USA
Listed on 2026-09-30
Listing for:
Wintermeyer Ventures
Full Time
position Listed on 2026-09-30
Job specializations:
-
IT/Tech
Data Engineering
Job Description & How to Apply Below
About the Role
This is a hands-on engineering role focused on building and operating large-scale web crawlers and data acquisition systems that power dataset creation for frontier AI model training. You will work as an extension of AI research teams, helping to collect, clean, and curate web data at a scale that few organizations can match. The work directly shapes the pretraining and inference pipelines used by some of the most advanced AI labs in the world.
WhatYou'll Do
- Build and maintain large-scale web crawlers across diverse domains including social media, travel, and multi-language sites.
- Design high-throughput, fault-tolerant data collection systems capable of handling millions to billions of URLs per day.
- Navigate anti-bot systems, rate limits, and JavaScript-heavy sites, finding creative solutions when standard protocols fall short.
- Develop pipelines for cleaning, deduplication, filtering, and normalization of web data at TB to PB scale.
- Construct and maintain datasets for research and model training in close collaboration with research teams.
- Monitor crawl performance, coverage, and data quality, iterating quickly as web environments change.
- Optimize infrastructure for cost, latency, and reliability across cloud and bare-metal environments.
- 3 or more years building and operating web crawlers at scale (2 or more years considered for candidates with a PhD).
- Proficiency in one or more of:
Go, Rust, Python, Java, or C++. - Experience running data pipelines at TB or greater scale.
- Deep knowledge of HTTP, networking, and browser behavior.
- Hands‑on experience with distributed systems or parallel processing.
- Experience with headless browsers such as Playwright, Puppeteer, or Chrome Dev Tools Protocol.
- Familiarity with proxy systems, IP rotation, or request orchestration.
- Experience with data quality evaluation, scoring, or benchmarking at scale.
- Experience running crawling or data workloads on cloud platforms (AWS, GCP) or bare-metal infrastructure.
- Background in NLP pipelines, ML dataset curation, or AI lab work is a strong plus.
Salary range: $160,000 to $250,000 USD annually. Visa sponsorship is not available for this role.
LocationThis role is fully remote. The primary location is Los Angeles, CA, United States, though candidates based in other major US cities are welcome.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×