scribeocr/scribeocr
Web interface for recognizing text, proofreading OCR, and creating fully-digitized documents.
What it solves
Scribe OCR is designed to solve the problem of inaccurate or poorly aligned OCR (Optical Character Recognition) data. While traditional OCR tools often produce "invisible text" layers that are roughly positioned over images, Scribe OCR focuses on high-precision alignment and efficient proofreading to move OCR accuracy from 98% to 100%.
How it works
The application runs entirely in the browser, ensuring no data is sent to remote servers. It uses the Scribe.js library for text recognition. To make errors easier to spot, the tool generates a custom font for each document, optimizing the alignment between the original scan and the overlay text. It provides a specific Proofreading Mode where editable text is precisely layered over source images, highlighting low-confidence characters in red to guide the user.
Who it’s for
It is intended for users who need to create fully digitized, ebook-style PDFs that replicate the original document's layout without the image background, as well as those who need to create accurate searchable PDFs or proofread existing OCR data (including Tesseract HOCR files).
Highlights
- Browser-based processing: All recognition and processing happens locally in the browser.
- Font Optimization: Generates custom fonts per document to improve text alignment and error detection.
- Ebook Mode: Produces text-native PDFs with small file sizes that faithfully replicate the original layout.
- Flexible Export: Supports both ebook-style PDFs and traditional invisible text-over-image PDFs.
- HOCR Support: Ability to edit and correct OCR data from other applications.
Related
- Project
- Project
- Project
- Project
- Project