Tomesphere Atlas: Interactive Mapping of 8.5 Million Research Papers
Overview of Tomesphere Atlas
Tomesphere Atlas is an interactive, high-density visualization tool designed to map millions of research papers. The system allows users to explore the vast landscape of academic research across multiple disciplines, providing a visual representation of how different fields of study evolve and cluster.
Key Features and Visualization Capabilities
Tomesphere Atlas provides several layers of visualization to help users navigate the research landscape. Each dot on the map represents a single research paper, where colors are assigned based on the primary research field.
- Field-Based Coloring: Users can color the map by Field, Year, or Citations.
- Density Mode: Users can actually see "where the papers are" through a density mode that includes a heatmap (a "soft watercolor cloud") that indicates where specific fields are densest.
- Field Lens: Users can isolate specific fields to see their distribution. The current view shows a strengths in Computer Science (48%), Physics (18%), and Mathematics (14%).
- Filtering and Tiers: To manage the data volume, the system provides tiers of data resolution:
- Top 25k recent papers (1 MB)
- Top 100k papers by citation (100k, 3 MB)
- Top 500k papers by citation (500k, 8 MB)
- All 3M papers (45 MB)
Data Scope and Temporal Analysis
The atlas maps a significant volume of academic work, including data from arXiv Life Sciences & Medicine. The temporal range of the papers indexed in the same view extends from 2007 to 2026, which includes projected or early-access papers.
The system allows users to analyze publication volume trends. For example, in a specific view, the year 2007 saw 10,151 papers, while the volume in 2020–2024 was roughly steady compared to 2015–2019 (with a ratio of 2020-2024 volume vs 2015-2019 volume being approximately 0.84).
Technical Infrastructure and Accessibility
Tomesphere Atlas is open and transparent regarding its technical implementation. The project provides access to the same assets used to build the map:
- GitHub repository for the source code.
- HuggingFace model and dataset for the the embeddings and the data used to generate the map.
- HuggingFace space for interactive exploration.
Users can also find papers via a search function (⌘K) and a dedicated Chrome extension to integrate the research discovery process into their browsing experience.
Sources
Related
- Project
- Dispatch
- Project
- Dispatch
- Project