LLMs Can't Jump: Analyzing the Limits of AI Scientific Discovery
The Core Thesis: LLMs Lack the Capacity for Foundational "Jumps"
In the position paper "LLMs Can't Jump," author Tom Zahavy argues that Large Language Models (LLMs) are structurally incapable of creating new foundational axioms or making the intuitive leaps required for major scientific breakthroughs. Using Albert Einstein's formulation of General Relativity as a primary case study, the paper posits that LLMs cannot perform the kind of "abductive reasoning" necessary to invent entirely new frameworks when observational data is scarce.
Zahavy suggests that the "jump" Einstein made—specifically the equivalence principle—was not a result of processing existing linguistic data, but was grounded in physical intuition and thought experiments based on sensory experience. The central claim is that because LLMs are trained on language (a lossy encoding of human experience) rather than direct interaction with the physical world, they lack the necessary grounding to simulate the world and run the internal experiments required for foundational invention.
Clarifications from the Author
Following the paper's circulation, Tom Zahavy provided critical context to clarify the scope and intent of his work:
- Personal Position, Not Corporate Policy: Zahavy emphasized that the paper is a personal position paper and does not represent the official view of Google DeepMind.
- Not a "Dead End" Argument: He clarified that the paper is not intended to argue that LLMs can never make scientific discoveries. He noted his own work on AlphaProof (the first AI system to win an IMO medal) as evidence that AI can and will continue to make amazing discoveries.
- Focus on a Specific Capability: The paper specifically explores what it would take for an AI to make a jump similar to the invention of General Relativity, rather than claiming AI is incapable of all forms of scientific progress.
Technical Critiques and Counter-Arguments
The thesis that LLMs cannot "jump" has met with significant skepticism from the technical community, focusing on three main areas:
1. Lack of Quantitative Evidence
Critics argue that the paper is a rhetorical position rather than an empirical study. One prominent critique suggests that the claim "LLMs can't jump" is not backed by quantitative evidence and lacks a rigorous definition of what constitutes a "jump." A proposed method for testing this would be to use an LLM with a 2025 knowledge cut-off and attempt to re-derive scientific results published in 2026 with minimal information.
2. The Role of Embodiment and Sensory Experience
While the paper argues that physical sensation and embodied simulation are necessary for discovery, some argue this is an over-rotation on the General Relativity analogy. For example, the quantization of energy in quantum mechanics was not discovered through physical sensation, but through the mathematical realization that quantization solved the black-body radiation spectrum problem. This suggests that "jumps" can occur in abstract domains without sensory grounding.
3. The "Moving Goalpost" Phenomenon
Several observers note that the benchmark for AI capability is constantly shifting. As LLMs destroy existing benchmarks, the definition of intelligence is moved to the next unreachable peak—in this case, the "intuitive jump." This perspective suggests that current limitations may be a function of the architecture or the "harness" (the environment) rather than a fundamental structural impossibility.
Proposed Solutions and Alternative Perspectives
To overcome the perceived limitation of "jumping," several theoretical paths have been proposed:
- World Models and Multimodal Grounding: The paper suggests that world models are the solution. However, critics point out that current multimodal fusion (integrating image/audio/video) has not yet yielded a general step-change in reasoning capabilities.
- Temperature Modulation: Some suggest that modulating LLM temperature—using high temperature for idea generation and low temperature for critique—could mimic the human process of "hallucinating" a possibility and then applying rigor to validate it.
- Environmental Feedback Loops: Integrating LLMs into physical world feedback loops (such as in automated chemistry research) may provide the grounding the author claims is missing.
- Architectural Changes: There is a suggestion that LLMs are fundamentally probabilistic and lack a mechanism for "orthogonal directional changes" in latent space, which would be necessary for intuitive leaps or humor. This would require a new architectural component specifically designed for making "left turns."
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch