Skip to main content

LangChain integration

Use Spider as a LangChain document loader. Crawled pages go straight into a chain or a RAG pipeline.

Install

Install LangChain and the Spider client.

Install the Spider client using pip

pip install spider_client

Put your API key in SPIDER_API_KEY.

Usage

Load documents from a URL with SpiderLoader.

Scrape with LangChain in Python

from langchain_community.document_loaders import SpiderLoader

loader = SpiderLoader(
    api_key="YOUR_API_KEY",
    url="https://spider.cloud",
    mode="scrape",
    # params={} # Optional parameters
)

data = loader.load()
print(data)

Modes

mode decides how much Spider collects.

ModeWhat it does
scrapeScrape one URL. The default.
crawlCrawl the site, following links to every page.

Custom parameters

Pass a params dictionary to change the request. The defaults are return_format: "markdown" and metadata: True. The API reference lists every parameter.