Scrapy

Crawl sites in Python with queues, parsing and pipelines. The current test measures single-page HTTP downloads, not crawling scale.

Recent field test
Source repositoryProject website

29 September 2026 UTC · frozen public-30-v1 · no proxy

One single-page Spider on all 30 targets

11 of 30 bodies met unchanged checks. Scrapy 2.19.0 / Twisted 26.4.0 used one fresh worker per URL with the stock HTTP/1.1 downloader, an identified Scrapy User-Agent, cookies and robots off, no proxy or JavaScript, and at most three ordinary redirects. Scrapy retry middleware was off; stock connection recovery was not separately instrumented. Inspect all 270 records across nine tools, including empty HTTP 200 checked text, a rejection missed by frozen phrases and the redirect-limit error. This tests the downloader, not Scrapy’s crawling or scheduling strengths. Raw captures stay private.

Compare all 30 targets across nine tools

Published test history

The 30-site single-page downloader comparison above is the current evidence. No May 2026 crawl run is published for this project.

What it can do

  • Crawl or queue pages
  • Fetch HTML
  • Extract fields

Capabilities are catalog notes, not outcomes from the benchmark above.

Project record

Type
Crawl many pages
Ecosystem
python
Repo check
2026-05-23
License
BSD-3-Clause

Repository metadata is a dated snapshot, not a live health guarantee.

See how we validate pages, including what a pass does and does not mean.