Skip to main content

Measure everything. Publish the method.

Spider Research is the lab layer of Spider Cloud: open benchmarks and field notes on web data quality for AI. Every number links to a method; every method links to a repo.

Stealth Bench V1 71 tasks · pass or block
Spider 93.9% best
Browser Use
81.4%
Anchor
76.8%
Onkernel
67.1%
Browserless
54%
Steel
47.2%
Browserbase
41.9%
Hyperbrowser
39.9%
Higher is better Reproduce

The record.

Each entry is a dated snapshot of a maintained measurement. Where a repo exists, the numbers reproduce from a clone.

Same tasks, same scoring, every provider.

A leaderboard is only as honest as its instrument. Every Spider benchmark holds three invariants:

01 Same tasks

Every provider runs the same 71-task list against the same pages, measuring completeness under anti-bot pressure. Each row comes from its own run, not one shared session.

02 Same scoring

A task passes when the page comes back and fails when the response matches one of the block patterns. Every provider gets the same check, and no model sits in the loop.

03 Open repo

We publish the task list and the scoring, and the runs we cite stay on this page. Clone spider-rs/benchmark and rerun the numbers yourself.

Method · spider-rs/benchmark ↗

On the bench.

  • Quality Bench v1

    Extraction fidelity, HTML against ground-truth markdown.

  • Dataset drop

    The 1,000-URL evaluation set behind the benchmark.

  • Silk structure eval

    Structure-conformance scoring for Silk output.

Judge the output yourself.

Post a URL, read what comes back. A failed request bills $0.