Public catalog
scraping
51 shown. Search ranks against problem language; browsing defaults to stronger public signals.
Scrapling
turnstile captcha blocking scraper
Verified 2026-08-10 · 68,957 stars
cloudflared
use cloudflared for scraping
Verified 2026-08-10 · 4,826 stars
HTTPie
debug HTTP APIs from the shell with readable JSON responses
Verified 2026-08-10 · 4,595 stars
lightpanda
use lightpanda for scraping
Verified 2026-08-10 · 4,292 stars
proxy-pool
rotating residential proxies
Verified 2026-08-10 · 4,080 stars
cheerio
fast HTML parsing in node without a browser
Verified 2026-08-10 · 3,221 stars
httpx / requests
async HTTP client for scrapers with timeouts and HTTP/2
Verified 2026-08-10 · 3,038 stars
Trafilatura
extract main article text and metadata from news HTML for RAG
Verified 2026-08-10 · 2,538 stars
headroom
use headroom for scraping
Verified 2026-08-10 · 1,814 stars
newspaper3k
extract news article title body and authors from article URLs
Verified 2026-08-10 · 1,713 stars
botasaurus
all-in-one anti-bot web scraping framework
Verified 2026-08-10 · 1,300 stars
requests-html
requests-like session that can render JS with pyppeteer for simple pages
Verified 2026-08-10 · 926 stars
firecrawl-mcp
Firecrawl as an MCP tool for agents to crawl URLs
Verified 2026-08-10 · 493 stars
lxml
parse HTML/XML at scale with lxml
Verified 2026-08-10 · 401 stars
selectolax
very fast HTML parser for large scrape corpora
Verified 2026-08-10 · 351 stars
cli-printing-press
CLI to print and package agent skill packs
Verified 2026-08-10 · 339 stars
grant
use grant for scraping
Verified 2026-08-10 · 320 stars
Splash
render JavaScript pages via a lightweight HTTP API for scrapy
Verified 2026-08-10 · 297 stars
Parsel
CSS and XPath extraction on HTML responses in scrapy-style code
Verified 2026-08-10 · 251 stars
got-scraping
HTTP client tuned for scraping with browser-like headers
Verified 2026-08-10 · 243 stars
geo-exit-ip-checklist
use geo-exit-ip-checklist for scraping
Verified 2026-08-10 · 213 stars
proxy-chain
programmable local proxy chain
Verified 2026-08-10 · 213 stars
html5lib
parse broken real-world HTML with a standards HTML5 parser
Verified 2026-08-10 · 117 stars
Beautiful Soup
parse static HTML into a navigable tree for scrapers
Verified 2026-08-10
Browserbase
hosted headless browsers for scraping or agents
Verified 2026-08-10
Camoufox
turnstile captcha blocking scraper
Verified 2026-08-10
cloudscraper
python client that bypasses cloudflare anti-bot
Verified 2026-08-10
Crawl4AI
LLM-friendly web crawling and extraction
Verified 2026-08-10
Crawlab
You run dozens of spiders in different languages and frameworks (Scrapy, Puppeteer, Selenium, plain scripts) and need one dashboard to deploy, schedule, and monitor them instead of cron plus ssh on each server.
Verified 2026-08-10
Crawlee
production crawlers with browser and HTTP queues
Verified 2026-08-10
Crawlee for Python
Your Python scraper keeps dying on transient errors and blocks and you need built-in retries, proxy rotation, and session management
Verified 2026-08-10
curl-impersonate
TLS fingerprint blocks plain curl or httpx
Verified 2026-08-10
fingerprint-suite
inject realistic browser fingerprints into playwright
Verified 2026-08-10
Firecrawl
API crawl that returns markdown from any URL
Verified 2026-08-10
FlareSolverr
self-hosted proxy that solves Cloudflare challenges
Verified 2026-08-10
Google Maps Scraper
You need an authorized CSV or JSON inventory of publicly listed businesses in a region for local-market, directory-quality, or coverage research
Verified 2026-08-10
GPT Crawler
I need to turn documentation pages from one site into a local JSON knowledge file.
Verified 2026-08-10
Jina Reader
turn a public web page into clean Markdown with one HTTP request
Verified 2026-08-10
MechanicalSoup
automate form login and multi-page HTML flows without a full browser
Verified 2026-08-10
Official API first
use an official structured API before scraping HTML
Verified 2026-08-10
OpenSERP
Your agent needs live Google or Bing results as structured JSON but a paid SERP API is too expensive per call
Verified 2026-08-10
Oxylabs AI-Crawler
I need to crawl a public domain and extract pages relevant to a natural-language research request.
Verified 2026-08-10
Oxylabs AI-Scraper
I need product fields from one public page as JSON but do not want to maintain CSS or XPath selectors.
Verified 2026-08-10
Oxylabs Google AI Mode Scraper API
You need to monitor how your brand or content is cited in Google AI Mode answers across different countries for SEO or GEO analysis
Verified 2026-08-10
PriceBuddy
You keep manually rechecking store pages for an item and want scheduled tracking with an alert when the price drops below your target
Verified 2026-08-10
Proxifly Free Proxy List
You are testing your own service's geolocation behavior and need a zero-cost pool of country-tagged exit IPs sorted by protocol
Verified 2026-08-10
proxy-agents
HttpsProxyAgent SocksProxyAgent
Verified 2026-08-10
ScrapeGraphAI
Your CSS or XPath scrapers keep breaking every time target sites change their layout and you want extraction driven by a natural-language prompt instead
Verified 2026-08-10
Scrapfly Anti-bot Detector
Your scraper suddenly returns challenge pages and you need to identify which anti-bot vendor (Cloudflare, Akamai, DataDome, PerimeterX, Kasada) protects the site
Verified 2026-08-10
Scrapy
large scale spider crawls with pipelines
Verified 2026-08-10
WaterCrawl
self-host asynchronous website crawling with a queue and API
Verified 2026-08-10