PDFMathTranslate/PDFMathTranslate
[EMNLP 2025 Demo] PDF scientific paper translation with preserved formats - 基于 AI 完整保留排版的 PDF 文档全文双语翻译,支持 Google/DeepL/Ollama/OpenAI 等服务,提供 CLI/GUI/MCP/Docker/Zotero
What it solves
It solves the problem of translating scientific PDF documents while preserving their original layout. Traditional translation often destroys the formatting of complex documents, but this tool ensures that formulas, charts, tables of contents, and annotations remain intact during the translation process.
How it works
The system uses a combination of precise layout detection (via DocLayout-YOLO) and large language models (LLMs) to identify and translate text while maintaining the document's structure. It supports multiple translation services (such as Google, DeepL, and OpenAI) and can be deployed as a command-line tool, a web-based interactive interface, or via Docker. It can generate either a mono-lingual translated document or a bilingual version.
Who it’s for
It is designed for researchers, students, and professionals who need to read scientific papers and technical documents in their native language without losing the critical visual context of the original layout.
Highlights
- Layout Preservation: Keeps formulas, charts, and annotations in their original positions.
- Flexible Deployment: Available as a CLI, GUI, Docker image, and Zotero plugin.
- Service Agnostic: Supports various translation backends including Google, DeepL, and MiniMax.
- Bilingual Output: Ability to generate documents with both source and target languages.
Written about in
Related
- Project
- Project
- Project
- Project
- Project