ICIJ/datashare
A self‑hosted search engine for documents
What it solves
Datashare provides a secure, self-hosted way for journalists and researchers to search and analyze massive amounts of heterogeneous data. It eliminates the need for external cloud services when handling sensitive leaked documents, allowing users to maintain full control over their material while performing complex searches across various file types.
How it works
The platform ingests diverse file formats (including PDFs, emails, spreadsheets, and images), extracts text using OCR for scans, and enriches the data with metadata and named-entity extraction. It indexes this information using Elasticsearch and stores it in PostgreSQL, exposing the results through a REST API and a dedicated search user interface.
Who it’s for
It is designed for investigative journalists, newsrooms, and researchers who need to process and query large-scale document leaks or archives securely.
Highlights
- Multi-format support: Indexes PDFs, emails, office documents, images, and archives.
- Automated enrichment: Automatically detects named entities such as people, organizations, and locations.
- Integrated OCR: Converts visual text in scans and images into searchable text.
- Advanced Querying: Supports boolean, wildcard, and fuzzy queries combined with faceted filters.
- Collaborative Tools: Includes a team/server mode for shared tags and recommendations.
Related
- Project
- Project
- Project
- Project
- Project