Crawlee

Crawl pages with queues, retries and browser options. The current test measures one-page CheerioCrawler HTTP downloads, not crawling scale.

Recent field test
Source repositoryProject website

29 September 2026 UTC · frozen public-30-v1 · no proxy

Crawlee HTTP on all 30 targets

13 of 30 source bodies met unchanged checks. This JavaScript Crawlee CheerioCrawler 3.18.2 pass used stock got-scraping, generated browser-like headers, three normal redirects and fresh sequential workers. Cookies, robots requests and crawler/session/block/got retries were disabled. No browser, JavaScript page execution, proxy or link enqueueing. Inspect all 330 dated records, including empty HTTP 200 checked text and Target’s HTTP 429 plus MIME error. This one-page test does not measure crawl scale, browser modes or the Python package, or establish a global winner. Raw bodies and logs stay private.

Compare all 30 targets across eleven tools

Published test history

The 30-site single-page HTTP comparison above is the current evidence. No May 2026 crawl run is published for this project.

What it can do

  • Keep a browser session
  • Crawl or queue pages
  • Render JavaScript
  • Take screenshots
  • Fetch HTML
  • Extract fields

Capabilities are catalog notes, not outcomes from the benchmark above.

Project record

Type
Crawl many pages
Ecosystem
python, typescript
Repo check
2026-05-23
License
Apache-2.0

Repository metadata is a dated snapshot, not a live health guarantee.

See how we validate pages, including what a pass does and does not mean.