LangChain integration
Use Spider as a LangChain document loader. Crawled pages go straight into a chain or a RAG pipeline.
Usage
Load documents from a URL with SpiderLoader.
Scrape with LangChain in Python
from langchain_community.document_loaders import SpiderLoader
loader = SpiderLoader(
api_key="YOUR_API_KEY",
url="https://spider.cloud",
mode="scrape",
# params={} # Optional parameters
)
data = loader.load()
print(data)Modes
mode decides how much Spider collects.
| Mode | What it does |
|---|---|
| scrape | Scrape one URL. The default. |
| crawl | Crawl the site, following links to every page. |
Custom parameters
Pass a params dictionary to change the request. The defaults are return_format: "markdown" and metadata: True. The API reference lists every parameter.