Document Management, OCR & PDF
14Essential tools for document digitization, management, and processing workflows
paperless-ngx
Self-hosted document management system
Stirling-PDF
Local web-based PDF toolbox
OCRmyPDF
Add OCR text layer to scanned PDFs
Paperless-AI
AI addon for paperless-ngx
paperless-gpt
ChatGPT integration for paperless-ngx
Tesseract
Industry-standard OCR engine
EasyOCR
Ready-to-use OCR with 80+ languages
ExifTool
Read/write metadata in files
markitdown
Convert documents to Markdown
pdfplumber
Plumb a PDF for detailed information about each char, rectangle, line, et cetera — and easily extract text and tables.
PyMuPDF
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
docTR
docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning.
Gotenberg
A developer-friendly API for converting many document formats into PDF files, and more!
WeasyPrint
The awesome document factory