Skip to main content
AI Studio  add-on for Spider.
Health
pubmed.ncbi.nlm.nih.gov Verified

PubMed Health Scraper

Extract biomedical literature citations, abstracts, author affiliations, and journal metadata from PubMed.

Get started Docs
target
pubmed.ncbi.nlm.nih.gov
success rate
99.9%
latency
~4ms
POST /fetch/pubmed.ncbi.nlm.nih.gov/
return_format
curl -X POST https://api.spider.cloud/fetch/pubmed.ncbi.nlm.nih.gov/ \
  -H "Authorization: Bearer $SPIDER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{}'
200 OK · cache hit · response shape
{
  "url": "https://pubmed.ncbi.nlm.nih.gov/",
  "status": 200,
  "data": {
    "article_title": "string",
    "authors": "string",
    "abstract": "string",
    "journal": "string",
    "doi": "string",
    "pmid": "string",
    "publication_date": "string",
    "keywords": "string"
  }
}
Quick start

Extract data in minutes.

Structured JSON from pubmed.ncbi.nlm.nih.gov with a single POST. Call it with no selectors and the model names the fields, or pass your own CSS to skip the AI.

pubmed-health-scraper.ts
import { SpiderBrowser } from "spider-browser";

const spider = new SpiderBrowser({
  apiKey: process.env.SPIDER_API_KEY!,
  stealth: 2,
});

await spider.connect();
const page = spider.page!;
await page.goto("https://pubmed.ncbi.nlm.nih.gov/?term=diabetes+treatment");

// No selectors, no schema. Spider reads the page and names the fields.
const data = await page.scrape();

console.log(data);
await spider.close();
ready to run · spider-browser · no selectors
Extraction

Fields you can pull.

Article titleAuthorsAbstractJournalDOIPMIDPublication dateKeywords
Content

Medical data extraction

Extract drug info, conditions, and health articles from pubmed.ncbi.nlm.nih.gov.

Parsing

Structured health data

Clean extraction of dosage, interactions, and clinical information.

Scale

Bulk research

Process thousands of medical pages for research and comparison datasets.

Related

More Health scrapers.

Start

Start scraping pubmed.ncbi.nlm.nih.gov.

Grab an API key and call the endpoint above. The first request resolves the config; every request after hits cache.