Measure everything. Publish the method.
Spider Research is the lab layer of Spider: open benchmarks and field notes on web data quality for AI. Every number links to a method; every method links to a repo.
The record.
Each entry is a dated snapshot of a maintained measurement. Where a repo exists, the numbers reproduce from a clone.
- Spider Browser Scores 85% on Browser Use's Stealth BenchmarkReproduce
Spider Browser scored 85% on Browser Use's open stealth benchmark, beating every other cloud browser tested against 80 anti-bot protected sites.
- Introducing Silk: Our Custom AI Model for Web Data Extraction
Spider runs Silk, a purpose-built extraction model that converts raw HTML into structured data and solves captchas on dedicated GPU infrastructure. No external API calls, no per-token billing, no data leaving our network.
- Spider Browser vs. Kernel vs. Browserbase: 999 URLs, 100% Pass RateReproduce
Kernel benchmarked cold start speed. We benchmarked what matters: reliability across 999 URLs, 254 domains, and 18 categories, with a 100% success rate and 2.5s median end-to-end latency.
- Crawl4AI vs Firecrawl vs Spider BenchmarkReproduce
Crawl4AI vs Firecrawl on 1,000 real URLs, with Spider in the mix: throughput, success rate, cost per 1K pages, RAG recall, licensing, and how to rerun it.
- Rust vs. Python for Web Scraping: Why We Rewrote Everything
The engineering story behind Spider's decision to abandon Python scrapers and rebuild from scratch in Rust. Concrete benchmarks, architecture decisions, and lessons learned.
- Scraping 1 Million Pages: What Actually Happens
An engineering log of crawling 1 million pages across 10,000 domains with Spider's cloud API. Throughput curves, failure modes, cost breakdown, and lessons learned.
- The True Cost of Web Scraping at Scale
A detailed cost breakdown of web scraping at 10K to 10M pages per month, comparing self-hosted Scrapy, Firecrawl, Apify, Crawl4AI, and Spider across infrastructure, proxies, engineering time, and total cost of ownership.
Same tasks, same judge, every provider.
A leaderboard is only as honest as its instrument. Every Spider benchmark holds three invariants:
Every provider runs the identical 71-task suite — completeness under anti-bot pressure, measured on the same pages in the same run.
One LLM judge scores every provider on the same rubric. The decimals in the leaderboard are judged pass rates, not rounding.
Methodology and raw results are public. Clone spider-rs/benchmark and the numbers rerun from your machine, not our word.
Method · spider-rs/benchmark ↗
On the bench.
- Quality Bench v1
Extraction fidelity — HTML against ground-truth markdown.
- Dataset drop
The 1,000-URL evaluation set behind the benchmark.
- Silk structure eval
Structure-conformance scoring for Silk output.
Judge the output yourself.
Post a URL, read what comes back. A failed request bills $0.