Public catalog
Browse tools
12 shown. Search ranks against problem language; browsing defaults to stronger public signals.
marker
convert PDFs to clean markdown for LLM pipelines
Verified 2026-08-10 ยท 37,362 stars
DeepSeek-OCR Dockerized API
Scanned or image-based PDFs have no text layer and you need structured markdown out of them
Verified 2026-08-10
docling
parse complex PDFs and docs into structured chunks for RAG
Verified 2026-08-10
dots.mocr
You need scanned or image-based multilingual documents parsed into structured layout JSON and markdown with correct reading order
Verified 2026-08-10
GLM-OCR
Scanned business documents with complex tables, seals, or code blocks come out garbled through conventional OCR and you need layout-aware markdown plus JSON output
Verified 2026-08-10
HunyuanOCR
I need local OCR and document parsing from a document image through a documented inference server.
Verified 2026-08-10
OCRmyPDF
add a searchable OCR text layer to scanned PDFs while preserving pages
Verified 2026-08-10
olmOCR
You have scanned or image-based PDFs with multi-column layouts, equations, tables, or handwriting and need Markdown or text in natural reading order.
Verified 2026-08-10
Open Notebook
open-source notebook LM style research notes
Verified 2026-08-10
pdf-inspector
A document pipeline sends every incoming PDF through a slow, expensive OCR service even though most are digitally generated with an intact text layer
Verified 2026-08-10
surya
use surya for evals testing
Verified 2026-08-10
unstructured
partition PDF HTML into structured elements
Verified 2026-08-10