Using Frontier LLMs to Decode Historical Cryptography and Trace Alchemical Knowledge

Frontier Large Language Models (LLMs) have evolved beyond simple research assistance to become capable of solving complex, tractable historical problems. By pairing domain experts with models such as GPT-6 (Sol/Astra) and Opus 5.5, researchers can now perform advanced cryptography, trace texts across translations, and synthesize links between disparate niche subfields that were previously invisible to human scholars.

Criteria for Tractable Historical Problems

AI models are most effective when applied to historical problems that meet specific "tractability" criteria. These problems typically involve tasks where a potential solution can be clearly proven or disproven, mirroring the way reasoning models have advanced in mathematics.

Key requirements for AI-driven historical breakthroughs include:

  • Digitized Data: The necessary primary sources must be fully digitized and accessible.
  • Expert Identification: Experts must have already defined the specific problems that need solving.
  • Multilingual Reasoning: The problem should leverage the model's ability to reason across different languages and disciplinary subfields.
  • Code Execution: The problem should be amenable to solutions involving bespoke code.

Case Studies in AI-Driven Historical Discovery

Recent applications of GPT-6 Astra and Opus 5.5 demonstrate the potential for these models to uncover new historical facts and verify existing ones.

Alchemical Links: Isaac Newton and Samuel Hartlib

Using Opus 5.5, researchers identified a compelling link between Isaac Newton and Samuel Hartlib. The model analyzed over 5,000 primary source files from the Hartlib archive and identified that both Newton and Hartlib used different anagrams to refer to "Hungarian vitriol," a key alchemical ingredient.

Opus 5.5 noted that while Newton used a true anagram ("Vltimorui"), Hartlib used a form of imperfect backwards writing ("Miloirtiua riciragnun"). This parallel, combined with matching quantities described in the texts, provides strong evidence that Newton drew upon Hartlib's manuscripts.

Historical Cryptography and Codebreaking

Frontier models have shown significant success in decrypting historical communications:

  • WWII Enigma Messages: GPT-6 Astra successfully broke a July 10, 1941, Enigma message that had resisted human decipherment for decades. The breakthrough occurred because the model noticed a specific note regarding radio message collections at the German Bundesarchiv that human experts had overlooked.
  • Emperor Charles V's Secret Code: Opus 5.5 partially deciphered 16th-century Spanish letters written in the secret code of Emperor Charles V. While these had been previously decrypted by scholars in 1916, the model's ability to independently reach the same conclusion served as a "gold standard" verification of its capabilities.
  • John Dee's Liber Loagaeth: In analyzing the coded manuscript of the occultist John Dee, GPT-6 Astra concluded that the text was largely nonsense syllables rather than a true code, though it did identify one specific reference to the angelic being "Bornogo."

The Bottleneck: Archival Access and Compute

Despite these capabilities, the primary limitation to historical AI research is not the model's reasoning, but the availability of data. Many historical manuscripts are digitized but remain behind privileged access walls, and the vast majority of premodern manuscripts remain undigitized.

To accelerate breakthroughs, the author proposes three interventions:

  1. Open Digitization: Collaborative efforts to make unavailable historical manuscripts freely accessible online.
  2. Compute Grants: Providing historians with free API access and high-compute environments to deploy hundreds of agents against active historical problems.
  3. Defining "Millennium Problems": Establishing a public list of solvable historical mysteries (e.g., the Voynich manuscript or Linear A) to focus AI efforts.

Community Insights and Counterpoints

Discussion among researchers and practitioners highlights both the promise and the risks of this approach:

"The prose is usually the least reliable part of the record... check the claim against the thing it points at rather than against another summary of it."

Some users noted the utility of AI in genealogical research to catch errors in widely shared family trees where secondary sources have blindly repeated a mistake. Others expressed skepticism regarding the interpretation of deeply metaphoric alchemical texts, suggesting that hermetic meanings may never be fully expressible in words or reducible to LLM reasoning.

Future Research Directions

The potential for AI in history extends beyond cryptography into provenance, the detection of historical plagiarism, and the reconstruction of ancient algorithms. For example, GPT-6 Pro is currently being used to reconstruct the solar eclipse algorithms used by the 17th-century Sanskrit astronomer Bhāskara in the Karaṇakesarī.

Sources

Related