Level

Thinking Machines Lab

Research Engineer, Web Crawling

AI in this role

Research engineer building and scaling internet-scale web crawling, deduplication, and ingestion pipelines to source pretraining data for AI models.

pythongorust
web-crawlingdistributed-systemsdata-engineeringscraping
About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Software Engineer to build and own our web-crawling systems, from distributed collection at internet scale through filtering, deduplication, and deciding what data we keep.

The ideal candidate has built and scaled a web crawler or large-scale data-acquisition systems. In this role, you'll write and own production systems: the crawler itself, the infrastructure that runs it at scale, and the pipelines that turn raw crawls into usable pretraining data. You'll work closely with our pretraining and data teams to understand what's actually moving model quality, but this is fundamentally an engineering role, not a research one.

What You'll Do
  • Design and scale the web crawler and ingestion infrastructure that sources Inkling's pretraining data

  • Build pipelines for large-scale extraction, deduplication, and data quality filtering

  • Build specialized crawlers for high-value or hard-to-reach data sources

  • Work with the pretraining team to understand how changes in crawled data affect model performance

  • Improve the reliability and efficiency of crawling and ingestion infrastructure at petabyte scale

  • Help set technical direction for this area as it grows, and bring other engineers up to speed on what you've learned

Skills & Qualifications

Minimum Qualifications

  • 8+ years designing, building, and scaling web crawlers, scrapers, or large-scale distributed data-acquisition systems

  • A track record of owning crawler or data-acquisition infrastructure at internet scale

  • Strong software engineering skills in a language such as Python, Go, or Rust, with real experience in distributed systems

  • Working knowledge of the practical and legal considerations of large-scale web data collection (robots.txt, rate limiting, licensing)

Preferred Qualifications

  • Experience applying machine learning to crawl selection, extraction, or data quality classification at internet scale

  • Experience setting technical direction for a crawling, data acquisition, or search infrastructure team, whether or not that was your formal title

  • Experience designing systems for petabyte-scale storage and processing

  • Track record of open-source contributions to crawling, scraping, or data infrastructure tools

  • Background at a search engine (crawling, indexing) or a frontier AI lab's data acquisition team

Logistics
  • Location: This role is based in San Francisco, CA.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000-$475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Thinking Machines Lab will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the California Fair Chance Act, the San Francisco Fair Chance Ordinance, and any other applicable state or local fair chance ordinance or law.

How we rate this

Research Engineer, Web Crawling at Thinking Machines Lab rates 80 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.

Classification

Builds AI. The job is building AI systems.

  1. ●●●● Builds AI80 to 100
  2. ●●●○ Works on AI60 to 79
  3. ●●○○ Uses AI40 to 59
  4. ●○○○ Little AI0 to 39

Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

Web CrawlingDistributed SystemsData EngineeringScrapingPythonGoRust

Questions you could be asked

  1. Tell me about a project where web crawling was part of your work. What did you do?
  2. Tell me about a project where distributed systems was part of your work. What did you do?
  3. Tell me about a project where data engineering was part of your work. What did you do?
  4. Tell me about a project where scraping was part of your work. What did you do?
  5. Walk me through how you've used Python in your day-to-day work.

Adapt your resume

  • List these exact terms on your resume: Web Crawling, Distributed Systems, Data Engineering, Scraping, and Python. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want an expert to read your CV for this job?

Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.

Get a free CV review

Get new AI jobs (Builds AI ●●●●) by email

One email a week with the new AI jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.

Free. One email a week. Unsubscribe in one click.

Similar roles

Research roles that build AI, at other companies.

What kind of AI work fits you?

Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.

Find my next step

More jobs at Thinking Machines Lab

Related searches

Same AI level