Same crawler. Less ops.
The core crawling engine is identical. What changes is everything around it: who manages the proxies, the browsers, the scaling, and the retries.
# you handle everything cargo add spider # set up proxy rotation # configure headless Chrome # manage server scaling # handle bans + retries # write markdown cleaning # build extraction pipeline
curl https://api.spider.cloud/crawl \ -H "Authorization: Bearer $KEY" \ -d '{ "url": "https://example.com", "limit": 100 }' → proxies, browsers, scaling handled
What changes when you go managed.
You size and manage your own fleet.
Cloud → Elastic. 10 pages or 10 million, same API call.
Bring your own. Rotate them yourself.
Cloud → Rotated for you. Blocked requests retry and cost nothing.
Headless Chrome plus community patches you keep up to date.
Cloud → A native browser on every request. Pages behind common protection services load completely.
You run headless Chrome or Firefox.
Cloud → Custom Rust browser. Faster, lighter, built for scraping.
html2md library output.
Cloud → Boilerplate stripped so the markdown is mostly the article, not the chrome.
Not included.
Cloud → Send a schema, get structured JSON. No parsers to maintain.
You want full control.
You have your own proxies, your own servers, and a team that can maintain the pipeline. Or you need to run on-prem for compliance reasons.
cargo add spiderYou want data, not infrastructure.
Sites are blocking you and you need proxy intelligence. Or you just want to ship faster and skip the ops work entirely.
curl https://api.spider.cloud/crawlBoth paths use the same Rust crawling engine. The open-source library is MIT-licensed and always will be. Hosted Spider is for when you want someone else to handle the hard parts.
Hosted Spider, free balance on signup.
No credit card required. Top up later when you're ready to scale.