Scrapy
Crawl sites in Python with queues, parsing and pipelines. The current test measures single-page HTTP downloads, not crawling scale.
29 September 2026 UTC · frozen public-30-v1 · no proxy
One single-page Spider on all 30 targets
11 of 30 bodies met unchanged checks. Scrapy 2.19.0 / Twisted 26.4.0 used one fresh worker per URL with the stock HTTP/1.1 downloader, an identified Scrapy User-Agent, cookies and robots off, no proxy or JavaScript, and at most three ordinary redirects. Scrapy retry middleware was off; stock connection recovery was not separately instrumented. Inspect all 270 records across nine tools, including empty HTTP 200 checked text, a rejection missed by frozen phrases and the redirect-limit error. This tests the downloader, not Scrapy’s crawling or scheduling strengths. Raw captures stay private.
Compare all 30 targets across nine toolsPublished test history
The 30-site single-page downloader comparison above is the current evidence. No May 2026 crawl run is published for this project.
What it can do
- Crawl or queue pages
- Fetch HTML
- Extract fields
Capabilities are catalog notes, not outcomes from the benchmark above.
Project record
- Type
- Crawl many pages
- Ecosystem
- python
- Repo check
- 2026-05-23
- License
- BSD-3-Clause
Repository metadata is a dated snapshot, not a live health guarantee.
See how we validate pages, including what a pass does and does not mean.