Public catalog
doc-parsing
24 shown. Search ranks against problem language; browsing defaults to stronger public signals.
marker
convert PDFs to clean markdown for LLM pipelines
Verified 2026-08-10 · 37,362 stars
tesseract
open-source OCR for images and scanned PDFs
Verified 2026-08-10 · 12,296 stars
mozilla-readability
extract main article text from messy HTML
Verified 2026-08-10 · 2,936 stars
DeepSeek-OCR 2
You need to convert scanned document images into structured markdown locally on your own GPU rather than through a cloud OCR API
Verified 2026-08-10
DeepSeek-OCR Client
You want to run DeepSeek-OCR on images locally through a point-and-click desktop app instead of writing Python inference scripts
Verified 2026-08-10
DeepSeek-OCR Dockerized API
Scanned or image-based PDFs have no text layer and you need structured markdown out of them
Verified 2026-08-10
docling
parse complex PDFs and docs into structured chunks for RAG
Verified 2026-08-10
dots.mocr
You need scanned or image-based multilingual documents parsed into structured layout JSON and markdown with correct reading order
Verified 2026-08-10
Easyocr
OCR images and screenshots with EasyOCR offline
Verified 2026-08-10
GLM-OCR
Scanned business documents with complex tables, seals, or code blocks come out garbled through conventional OCR and you need layout-aware markdown plus JSON output
Verified 2026-08-10
HunyuanOCR
I need local OCR and document parsing from a document image through a documented inference server.
Verified 2026-08-10
LiteParse
You must parse PDFs, Office documents, or scanned images to markdown or structured JSON without uploading them to any cloud service
Verified 2026-08-10
OCRmyPDF
add a searchable OCR text layer to scanned PDFs while preserving pages
Verified 2026-08-10
Ollama OCR
I need to extract text or structured output from an image or PDF with a local Ollama vision model.
Verified 2026-08-10
olmOCR
You have scanned or image-based PDFs with multi-column layouts, equations, tables, or handwriting and need Markdown or text in natural reading order.
Verified 2026-08-10
paddleocr
PaddleOCR for multilingual OCR pipelines
Verified 2026-08-10
pdf-inspector
A document pipeline sends every incoming PDF through a slow, expensive OCR service even though most are digitally generated with an intact text layer
Verified 2026-08-10
Pdfplumber
extract PDF tables and text with pdfplumber
Verified 2026-08-10
pix2tex (LaTeX-OCR)
You have a screenshot or image of a printed math formula from a paper and need the corresponding LaTeX source instead of retyping it symbol by symbol.
Verified 2026-08-10
RapidOCR
Text must be extracted from images with Chinese and English content fully offline, without sending documents to a cloud OCR API
Verified 2026-08-10
surya
use surya for evals testing
Verified 2026-08-10
Umi-OCR
Hundreds of local images or screenshots need their text extracted in batch, fully offline
Verified 2026-08-10
unstructured
partition PDF HTML into structured elements
Verified 2026-08-10
WhatsApp Backup Reader
You exported a large WhatsApp chat as a .zip and the raw _chat.txt plus loose media files are impossible to actually browse or search
Verified 2026-08-10