ScrapingEvals

Scraping tools, tested against real protected sites.

No affiliate listicles, no vendor-supplied numbers. We run each tool against 21 real public sites — Amazon, Cloudflare-fronted stores, travel and real-estate portals — and save the request, the response, and the screenshot for every attempt. Then we write down what actually happened.

10 tools reviewed with real runs26 in the catalog21 live targets each
New · field note 001One G2 request moved through two failed ScrapeDrive tiers before Hyperdrive returned validated content.Inspect trace →

Tools tested · public-tough-v1 · proxy: none · 2026-05-24

ToolTypeVerdictPassedNotes
BotasaurusAnti-detect browserUsable0/21Best untuned score in the test — from its HTTP layer alone.
CamoufoxAnti-detect browserSharp edges0/21The highest ceiling we measured — if you pay the babysitting tax.
curl_cffiTLS-impersonation HTTPUsable0/21Chrome's TLS handshake without the browser.
HTTPXHTTP clientUsable0/21Modern sync/async HTTP for Python, same wall as Requests.
PatchrightBrowser automationSharp edges0/21Measured against stock Playwright: no difference (in this mode).
PlaywrightBrowser automationUsable0/21A rendering baseline, not a stealth tool.
Python RequestsHTTP clientUsable0/21The Python HTTP baseline — and our negative control.
ScraplingTLS-impersonation HTTPUsable0/21curl_cffi-class fetching with a scraping-first API.
SeleniumBaseBrowser automationSharp edges0/21Beats Playwright on Amazon, fails murkier everywhere else.
wreqTLS-impersonation HTTPSharp edges0/21Real fingerprint, young edges.

Pass counts are out of 21 and measure the tool and its network position together. With no proxy, a low score often reflects a single residential IP's reputation as much as the tool. That is why these are reviews, not a leaderboard — read the per-tool page for what each number means. Camoufox's row shows its strongest of three configurations.

How we test

  1. Run the tool in its natural shape against 21 real public pages.
  2. Judge each attempt with deterministic validators: expected status code plus required and forbidden content markers — not an LLM guess.
  3. Save the request config, response metadata, HTML, and screenshots as artifacts.
  4. Write an honest verdict: what worked, what failed, and when you'd reach for it.

What this isn't (yet)

These are open-source tools you run yourself. We have not published numbers for hosted scraping APIs (Zyte, Scrapfly, Bright Data, and the like): the adapters exist, but no real paid runs have been made, so there is nothing honest to show. Proxy-mode results are also still to come. We'd rather ship a small true map than a big fake one.