LLMs can’t jump (Zahavy, Google DeepMind, 2026)
tom-zahavy, Google DeepMind. Dated 27 January 2026, 10 pages, a position paper — the author’s word. T1: the PDF read directly, after the press version arrived first and got the claim wrong.
The argument
Einstein, in a letter to Maurice Solovine, described discovery as a cycle: from sense experience (E) an intuitive jump (J) to a system of axioms (A), then logical deduction to theorems, then verification back against experience. Zahavy maps three inference types onto it and asks which ones machines have.
- Induction — statistical pattern matching. Mastered.
- Deduction — formal derivation from established premises. Being conquered, and he cites alphaproof (Hubert et al., 2025) as the evidence.
- Abduction — generating a novel explanatory hypothesis for a singular phenomenon. Missing, and he argues structurally missing rather than merely absent.
The case study is General Relativity, treated as a computational problem. The target is Schmidhuber’s (2008) Theory of Compression Progress — discovery as the search for a program that compresses observations — which makes discovery a species of induction. If discovery were induction plus deduction, a sufficiently large model should be able to invent General Relativity given compute.
Zahavy’s counterexample is that GR was formulated where the data was scarce. There was no error signal in Newtonian observation demanding replacement; Einstein was driven by a conceptual inconsistency between action at a distance and field theories of electromagnetism. Compression has nothing to compress. Penrose, quoted in the paper: “Einstein was not just ‘noticing patterns’… He was uncovering profound mathematical structure that was already hidden in the very working of the world.”
What supplies the jump: manipulative abduction
Einstein’s “happiest thought” — an observer falling freely feels no gravitational field, and so may interpret himself as at rest — is the mechanism. Zahavy names it manipulative abduction (Magnani et al., 2009): hypotheses generated by embodied simulation, thinking by doing, actively intervening in a mental model rather than observing one.
He splits the thought experiment into two stages. An observation is imagined via simulation, then an explanation is derived by abduction. ARC-AGI tests only the second. Its grids are too sparse for induction and lack the instructions deduction needs, so a solver must make a logical leap — but the manipulative half, the physical sensation driving Einstein’s insight, isn’t in it. That’s the paper’s sharpest small claim, and it’s a criticism of a benchmark this hub takes seriously.
The concession is explicit and matters: an LLM could plausibly execute the deductive phase if handed Einstein’s premises. The 1913–1915 grind — Einstein and Grossmann searching geometric constraints, finding the Riemann curvature tensor and wrongly discarding it — “closely mirrors the capabilities of modern neuro-symbolic AI.” The claim is not that machines can’t do physics. It is that they can’t produce the axioms they then reason from.
The proposed bridge
Not “never.” The paper identifies translating simulation into formal axioms as the critical bottleneck and proposes physically consistent multimodal world models as the way across, with a distinction that carries the argument:
- Passive video prediction is not enough. Veo-class models show intuitive physics as “a byproduct of statistical correlation; they correctly generate a falling apple not because they model gravity, but because falling is the dominant continuation of unsupported object in their training distribution.”
- Action-controllable world models are the shift. Genie (Bruce et al., 2024) learns an action space permitting agentic intervention, which is the prerequisite for thinking by doing. “To replicate Einstein’s elevator thought experiment, an AI cannot merely watch a video of an elevator; it must possess the capacity for counterfactual intervention… It must be able to essentially take control of the simulation to conceptually cut the cable.”
The ambition is stated plainly: turn the abductive jump “from a mystical insight into a reproducible algorithmic process.” Also invoked: Harnad’s symbol grounding problem — LLMs as high-dimensional “Chinese Rooms” manipulating the language of physics without its referents — plus LeCun on world models and Li on spatial intelligence.
A scope limit the press version dropped entirely: the proposal is tailored to the physical sciences. In mathematics and computer science the “sense experience” is grounded differently — high-dimensional topology, or goals like generality and minimality — so “the necessity of the Abductive Jump remains universal, the nature of the simulation must be adapted to the ontology of the discipline.”
Why it lands here
Cluster E is where this spoke keeps the mechanization of reasoning, and alphaproof is its AI↔formal-methods bridge. This paper is the boundary condition for that programme, written by one of AlphaProof’s own co-authors: deduction is being mechanized and the axiom-generating step is not, and the reason offered is grounding rather than scale. It is also the strongest counter in the corpus to the implicit “more compute closes any gap,” made from inside the lab with the most to gain from that claim.
Related
tom-zahavy · abduction · alphaproof · llms-cant-jump-press · lean-theorem-prover · synthesis