Skip to main content
AI Studio  add-on for Spider.
sharechat.com · HTTP 200

Sharechat Scraper

Spider read sharechat.com in 5.8 s without a browser and returned 111 lines of clean markdown, including sections like "4. Feature Service" and "Step 1: Make db schema more “compact".

Get your free API key
Free balance on signup No card. Failed requests cost $0.
Response sharechat.com/blogs/artificial-intelligence/how-sharechat-built-a-scalable-cost-efficient-ml-feature-system.md markdown · 111 lines
We use ScyllaDB as our storage backend, which provides high throughput and low latency access to our pre-aggregated tiles. ScyllaDB's architecture is particularly well-suited for our use case as it can handle high write loads from our Flink jobs while simultaneously serving read requests from our Feature Service.### **4. Feature Service**The Feature Service is our serving layer that handles real-time feature requests from the recommendation system. When a feature request arrives, this service:* Checks the local cache if the same request has been already processed recently. Returns the result immediately if it did, if not:* Determines which tiles are required to compute the requested feature* Retrieves the relevant tiles from ScyllaDB* Performs final aggregations across tiles to compute the exact feature values* Returns the computed features to the calling service and caches the computed result# **Tiles optimizations**The initial database schema and tiling configuration led to scalability problems. Original schema mapped each entity into its own partition, with timestamp and feature name being ordered clustering columns.Tiles were computed for segments of one minute, 30 minutes and one day. The most popular requested aggregation ranges were 1 hour, 1 day, 7 days or 30 days, and the number of tiles required to be fetched were 70 per feature on average.If we do the math, it becomes obvious why this approach didn’t scale well. At the moment of testing, the system has around 8K rps for fetching the feed, with around 2K candidates being ranked:### **Step 1: Make db schema more “compact”**
Code · Fields · Cost · Run it keyless, no account

The same call, in code.

The capture above came back as markdown. These examples add a key, so you get browser rendering, proxies, and concurrency on sharechat.com.

sharechat-com-scraper.ts
import { SpiderBrowser } from "spider-browser";

const spider = new SpiderBrowser({
  apiKey: process.env.SPIDER_API_KEY!,
});

await spider.connect();
const page = spider.page!;
await page.goto("https://sharechat.com");

// No selectors, no schema. Spider reads the page and names the fields.
const data = await page.scrape();

console.log(data);
await spider.close();
ready to run · spider-browser, no selectors

Ready for volume? Get an API key →

Fields you can pull.

UsernamePost ContentLikesCommentsTimestamp

Spider names these from the page. The capture above came back as markdown; the same call with return_format: "json" returns them as keys.

What sharechat.com costs to scrape.

The capture above cost $0.000038 to fetch. Pricing is $1 per GB of pre-transformation bandwidth plus $0.001 per CPU minute, so a page like this one lands at a fraction of a cent. Failed requests are billed at $0.

  • Free balance on signup
  • No card required to test
  • Balance never expires
See the full pricing →

Run it keyless, no account

curl -X POST https://api.spider.cloud/scrape -H "Content-Type: application/json" -d '{"url": "https://sharechat.com/blogs/artificial-intelligence/how-sharechat-built-a-scalable-cost-efficient-ml-feature-system", "return_format": "markdown"}'

More Social scrapers.

Start scraping sharechat.com.

You already have the call. A key raises the rate limit and turns on browser rendering, proxies, and concurrency. Balance never expires, and top-ups go through secure checkout.