Skip to main content
AI Studio  add-on for Spider.
Spider Research

Measure everything. Publish the method.

Spider Research is the lab layer of Spider: open benchmarks and field notes on web data quality for AI. Every number links to a method; every method links to a repo.

Stealth Bench V1 71 tasks · LLM judge
Spider 84.5% best
Browser Use
81.4%
Anchor
76.8%
Onkernel
67.1%
Browserless
54%
Steel
47.2%
Browserbase
41.9%
Hyperbrowser
39.9%
Higher is better Reproduce
Publications

The record.

Each entry is a dated snapshot of a maintained measurement. Where a repo exists, the numbers reproduce from a clone.

Method

Same tasks, same judge, every provider.

A leaderboard is only as honest as its instrument. Every Spider benchmark holds three invariants:

01 Same tasks

Every provider runs the identical 71-task suite — completeness under anti-bot pressure, measured on the same pages in the same run.

02 Same judge

One LLM judge scores every provider on the same rubric. The decimals in the leaderboard are judged pass rates, not rounding.

03 Open repo

Methodology and raw results are public. Clone spider-rs/benchmark and the numbers rerun from your machine, not our word.

Method · spider-rs/benchmark ↗

Next

On the bench.

  • Quality Bench v1

    Extraction fidelity — HTML against ground-truth markdown.

  • Dataset drop

    The 1,000-URL evaluation set behind the benchmark.

  • Silk structure eval

    Structure-conformance scoring for Silk output.

Judge the output yourself.

Post a URL, read what comes back. A failed request bills $0.