UB-Mannheim/zotero-ocr
Zotero Plugin for OCR
What it solves
It enables users to perform Optical Character Recognition (OCR) on PDF files directly within Zotero, making non-searchable PDFs searchable and extracting text from images of documents.
How it works
The plugin integrates Tesseract OCR and pdftoppm (from Poppler tools) into Zotero. When a user selects a PDF and triggers the OCR process via the context menu, the tool processes the document and can generate a new PDF with a text layer, a note containing the recognized text, or HTML (hOCR) files.
Who it’s for
Researchers, students, and academics who use Zotero to manage their libraries and need to make scanned or image-based PDFs searchable and accessible.
Highlights
- Tesseract Integration: Uses the industry-standard Tesseract OCR engine for text recognition.
- Flexible Output: Can create new PDFs with recognized text, text-only notes, or hOCR files.
- Customizable Settings: Allows users to adjust output DPI, Tesseract Page Segmentation Mode (PSM), and language/script models.
- Attachment Management: Supports adding processed PDFs as either normal attachments or linked files.
Related
- Project
- Project
- Project
- Project
- Project