Research Crawling Engineer
- Location
- Los Angeles, San Francisco +1
- Workplace
- Remote
- Compensation
- $160k – $250k + equity
- Visa
- No Visa Sponsorship
About this role
What we're looking for:
We need someone with 4+ years of experience (2+ for candidates with PhDs) building web crawlers and large-scale data acquisition systems who has hands-on experience designing high-throughput, fault-tolerant pipelines. You should be comfortable operating at the boundary of scale and reliability in adversarial web environments and have a track record of processing millions to billions of URLs/day. Bonus points if you have experience with NLP pipelines, LLM pretraining data, or dataset curation for ML.
What you'll do:
Build and maintain large-scale web crawlers across diverse domains (social media, travel, multi-language sites) that power dataset creation for frontier AI labs
Design high-throughput, fault-tolerant systems for data collection handling millions to billions of URLs/day
Handle anti-bot systems, rate limits, and dynamic/JS-heavy sites — thinking creatively when standard protocols fail
Develop pipelines for cleaning, deduplication, filtering, and normalization of web data at TB–PB scale
Construct and maintain datasets for research and model training, collaborating directly with research teams to align data collection with modeling needs
Monitor crawl performance, coverage, and data quality; iterate quickly as web environments constantly change
Optimize infrastructure for cost, latency, and reliability across cloud or bare-metal environments
About the team:
We build infrastructure that delivers massive amounts of web data to the companies training the world's most powerful AI models. Frontier AI labs are our customers — you'll be working as an extension of their data teams on cutting-edge pre-training and inference model development. We own one of the largest repositories of public web data and have more resources available for working with data at scale than basically any other company. We're a lean, flat organization (~42 people) with no people managers — just builders pushing to expand what's possible for open web data and AI. We're cash flow positive and growing quickly.
What happens next
Skip the application pile. I get you in front of the people who decide.
Confirm the fit
A few questions to make sure this role is the right shape for you. Two minutes.
I pitch you to the company
I write the intro, send it to the founder, and handle the back-and-forth.
A meeting lands on your calendar
When the company wants to meet, I get the call on your calendar. You just show up.
Culture & values
The organization is flat by design with absolutely no people managers.
The CTO serves as the product manager across every workstream and every team member is a self-directed individual contributor.
Communication is paramount in a flat structure and proactive communication is non-negotiable.
Low ego, high output; no job is too big or too small and everyone contributes wherever needed.
The team values people who make everyone around them better, not just those who keep things moving.
The culture emphasizes building hard problems and doing work that matters, not comfort.
Self-starters who learn fast are sought after; if you don't know something, you research it and figure it out.
The team is fully remote and globally distributed, with no red tape and fast decision-making.
Teams collaborate across the organization rather than remaining siloed in narrow functions.
The environment prioritizes substance over ceremony and pushing the boundaries of what's possible with open web data and AI.
Know someone who'd be great for this?