MerchantryTidbits

Public catalog

scraping

51 shown. Search ranks against problem language; browsing defaults to stronger public signals.

scrapinglibraryfree

Scrapling

turnstile captcha blocking scraper

Verified 2026-08-10 · 68,957 stars

scrapingclifree

cloudflared

use cloudflared for scraping

Verified 2026-08-10 · 4,826 stars

scrapingclifree

HTTPie

debug HTTP APIs from the shell with readable JSON responses

Verified 2026-08-10 · 4,595 stars

scrapingclifree

lightpanda

use lightpanda for scraping

Verified 2026-08-10 · 4,292 stars

scrapingclifree

proxy-pool

rotating residential proxies

Verified 2026-08-10 · 4,080 stars

scrapinglibraryfree

cheerio

fast HTML parsing in node without a browser

Verified 2026-08-10 · 3,221 stars

scrapinglibraryfree

httpx / requests

async HTTP client for scrapers with timeouts and HTTP/2

Verified 2026-08-10 · 3,038 stars

scrapinglibraryfree

Trafilatura

extract main article text and metadata from news HTML for RAG

Verified 2026-08-10 · 2,538 stars

scrapinglibraryfree

headroom

use headroom for scraping

Verified 2026-08-10 · 1,814 stars

scrapinglibraryfree

newspaper3k

extract news article title body and authors from article URLs

Verified 2026-08-10 · 1,713 stars

scrapinglibraryfree

botasaurus

all-in-one anti-bot web scraping framework

Verified 2026-08-10 · 1,300 stars

scrapinglibraryfree

requests-html

requests-like session that can render JS with pyppeteer for simple pages

Verified 2026-08-10 · 926 stars

scrapingmcpapi_key

firecrawl-mcp

Firecrawl as an MCP tool for agents to crawl URLs

Verified 2026-08-10 · 493 stars

scrapinglibraryfree

lxml

parse HTML/XML at scale with lxml

Verified 2026-08-10 · 401 stars

scrapinglibraryfree

selectolax

very fast HTML parser for large scrape corpora

Verified 2026-08-10 · 351 stars

scrapingclifree

cli-printing-press

CLI to print and package agent skill packs

Verified 2026-08-10 · 339 stars

scrapinglibraryfree

grant

use grant for scraping

Verified 2026-08-10 · 320 stars

scrapingapifree

Splash

render JavaScript pages via a lightweight HTTP API for scrapy

Verified 2026-08-10 · 297 stars

scrapinglibraryfree

Parsel

CSS and XPath extraction on HTML responses in scrapy-style code

Verified 2026-08-10 · 251 stars

scrapinglibraryfree

got-scraping

HTTP client tuned for scraping with browser-like headers

Verified 2026-08-10 · 243 stars

scrapingpatternfree

geo-exit-ip-checklist

use geo-exit-ip-checklist for scraping

Verified 2026-08-10 · 213 stars

scrapinglibraryfree

proxy-chain

programmable local proxy chain

Verified 2026-08-10 · 213 stars

scrapinglibraryfree

html5lib

parse broken real-world HTML with a standards HTML5 parser

Verified 2026-08-10 · 117 stars

scrapinglibraryfree

Beautiful Soup

parse static HTML into a navigable tree for scrapers

Verified 2026-08-10

scrapingapipaid

Browserbase

hosted headless browsers for scraping or agents

Verified 2026-08-10

scrapinglibraryfree

Camoufox

turnstile captcha blocking scraper

Verified 2026-08-10

scrapinglibraryfree

cloudscraper

python client that bypasses cloudflare anti-bot

Verified 2026-08-10

scrapinglibraryfree

Crawl4AI

LLM-friendly web crawling and extraction

Verified 2026-08-10

scrapingappfree

Crawlab

You run dozens of spiders in different languages and frameworks (Scrapy, Puppeteer, Selenium, plain scripts) and need one dashboard to deploy, schedule, and monitor them instead of cron plus ssh on each server.

Verified 2026-08-10

scrapinglibraryfree

Crawlee

production crawlers with browser and HTTP queues

Verified 2026-08-10

scrapinglibraryfree

Crawlee for Python

Your Python scraper keeps dying on transient errors and blocks and you need built-in retries, proxy rotation, and session management

Verified 2026-08-10

scrapingclifree

curl-impersonate

TLS fingerprint blocks plain curl or httpx

Verified 2026-08-10

scrapinglibraryfree

fingerprint-suite

inject realistic browser fingerprints into playwright

Verified 2026-08-10

scrapingapiapi_key

Firecrawl

API crawl that returns markdown from any URL

Verified 2026-08-10

scrapingclifree

FlareSolverr

self-hosted proxy that solves Cloudflare challenges

Verified 2026-08-10

scrapingclifree

Google Maps Scraper

You need an authorized CSV or JSON inventory of publicly listed businesses in a region for local-market, directory-quality, or coverage research

Verified 2026-08-10

scrapingclifree

GPT Crawler

I need to turn documentation pages from one site into a local JSON knowledge file.

Verified 2026-08-10

scrapingapifree

Jina Reader

turn a public web page into clean Markdown with one HTTP request

Verified 2026-08-10

scrapinglibraryfree

MechanicalSoup

automate form login and multi-page HTML flows without a full browser

Verified 2026-08-10

scrapingpatternfree

Official API first

use an official structured API before scraping HTML

Verified 2026-08-10

scrapingclifree

OpenSERP

Your agent needs live Google or Bing results as structured JSON but a paid SERP API is too expensive per call

Verified 2026-08-10

scrapingapipaid

Oxylabs AI-Crawler

I need to crawl a public domain and extract pages relevant to a natural-language research request.

Verified 2026-08-10

scrapingapipaid

Oxylabs AI-Scraper

I need product fields from one public page as JSON but do not want to maintain CSS or XPath selectors.

Verified 2026-08-10

scrapingapipaid

Oxylabs Google AI Mode Scraper API

You need to monitor how your brand or content is cited in Google AI Mode answers across different countries for SEO or GEO analysis

Verified 2026-08-10

scrapingapplocal

PriceBuddy

You keep manually rechecking store pages for an item and want scheduled tracking with an alert when the price drops below your target

Verified 2026-08-10

scrapingapifree

Proxifly Free Proxy List

You are testing your own service's geolocation behavior and need a zero-cost pool of country-tagged exit IPs sorted by protocol

Verified 2026-08-10

scrapinglibraryfree

proxy-agents

HttpsProxyAgent SocksProxyAgent

Verified 2026-08-10

scrapinglibraryapi_key

ScrapeGraphAI

Your CSS or XPath scrapers keep breaking every time target sites change their layout and you want extraction driven by a natural-language prompt instead

Verified 2026-08-10

scrapingappfree

Scrapfly Anti-bot Detector

Your scraper suddenly returns challenge pages and you need to identify which anti-bot vendor (Cloudflare, Akamai, DataDome, PerimeterX, Kasada) protects the site

Verified 2026-08-10

scrapinglibraryfree

Scrapy

large scale spider crawls with pipelines

Verified 2026-08-10

scrapinglibraryfree

WaterCrawl

self-host asynchronous website crawling with a queue and API

Verified 2026-08-10