Yuliang-Liu/MultimodalOCR
On the Hidden Mystery of OCR in Large Multimodal Models (OCRBench)
What it solves
This repository provides a suite of benchmarks to evaluate how well Large Multimodal Models (LMMs) handle visual text. It addresses the lack of systematic evaluation for multilingual document parsing, real-world photographed documents, and complex visual text localization and reasoning tasks across diverse scripts and languages.
How it works
The project consists of three primary benchmarks:
- MDPBench: Focuses on multilingual document parsing across 17 languages and diverse scripts, testing models on both clean digital pages and real-world photographed documents.
- OCRBench v2: A large-scale bilingual benchmark that tests visual text localization and reasoning across 31 diverse scenarios (such as receipts, formulas, and street scenes) using 10,000 human-verified QA pairs.
- OCRBench: The original benchmark assessing core OCR capabilities, including text recognition, scene-text VQA, document VQA, key information extraction, and handwritten mathematical expression recognition.
Who it’s for
Researchers and developers building or evaluating Large Multimodal Models (LMMs) who need to measure their performance on optical character recognition (OCR) and document understanding tasks.
Highlights
- Multilingual Scope: MDPBench covers 17 languages, including low-resource languages and non-Latin scripts.
- Real-World Testing: Specifically evaluates performance on photographed documents versus digital ones.
- Comprehensive Scenarios: OCRBench v2 covers 31 different visual text scenarios.
- Human-Verified: All benchmarks utilize high-quality, human-verified annotations to ensure precise evaluation.
Related
- Project
- Project
- Project
- Project
- Dispatch