Google Scholar Scraper
Spider read scholar.google.com in 282 ms without a browser and returned 33 lines of clean markdown, including the section "Języki".
Nie można teraz wykonać tej operacji. Spróbuj ponownie później.Pokaż prace, których **autorem** jestnp. *"Maria Skłodowska"* lub *Mendelejew*np. *Świat Nauki* lub *Wiedza i Życie*)## Zapisano w Mojej biblioteceMój profilMoja bibliotekaLaboratoriumDowolny język Tylko język polski### Języki The same call, in code.
The capture above came back as markdown. These examples add a key, so you get browser rendering, proxies, and concurrency on scholar.google.com.
import { SpiderBrowser } from "spider-browser";
const spider = new SpiderBrowser({
apiKey: process.env.SPIDER_API_KEY!,
});
await spider.connect();
const page = spider.page!;
await page.goto("https://scholar.google.com/scholar?q=machine+learning+transformers");
// No selectors, no schema. Spider reads the page and names the fields.
const data = await page.scrape();
console.log(data);
await spider.close(); import { SpiderBrowser } from "spider-browser";
const spider = new SpiderBrowser({
apiKey: process.env.SPIDER_API_KEY!,
});
await spider.connect();
const page = spider.page!;
await page.goto("https://scholar.google.com/scholar?q=machine+learning+transformers");
const data = await page.evaluate(`(() => {
const papers = [];
document.querySelectorAll(".gs_r.gs_or.gs_scl").forEach(el => {
const title = el.querySelector(".gs_rt a")?.textContent?.trim();
const authors = el.querySelector(".gs_a")?.textContent?.trim();
const snippet = el.querySelector(".gs_rs")?.textContent?.trim();
const citations = el.querySelector(".gs_fl a")?.textContent?.trim();
const link = el.querySelector(".gs_rt a")?.getAttribute("href");
if (title) papers.push({ title, authors, snippet, citations, link });
});
return JSON.stringify({ total: papers.length, papers: papers.slice(0, 10) });
})()`);
console.log(JSON.parse(data));
await spider.close(); Ready for volume? Get an API key →
Fields you can pull.
Spider names these from the page. The capture above came back as markdown; the same
call with return_format: "json" returns them as keys.
What scholar.google.com costs to scrape.
The capture above cost $0.000213 to fetch. Pricing is $1 per GB of pre-transformation bandwidth plus $0.001 per CPU minute, so a page like this one lands at a fraction of a cent. Failed requests are billed at $0.
- Free balance on signup
- No card required to test
- Balance never expires
Run it keyless, no account
More Education scrapers.
Coursera Scraper
Extract course listings, instructor data, ratings, and enrollment info from Coursera.
Udemy Scraper
Extract course listings, pricing, instructor reviews, and curriculum data from Udemy.
Amazon Books Scraper
Extract bestseller book data, ratings, pricing, and author info from Amazon Books.
Start scraping scholar.google.com.
You already have the call. A key raises the rate limit and turns on browser rendering, proxies, and concurrency. Balance never expires, and top-ups go through secure checkout.