HN Debrief

Position: LLMs Can't Jump

  • AI
  • Science
  • Research
  • Machine Learning

The paper is a position piece, not an empirical result. It uses Einstein’s path to general relativity as a case study and argues that LLMs are good at induction and increasingly good at deduction, but still lack abduction, meaning the ability to form new explanatory hypotheses from sparse evidence. The core claim is that Einstein’s leap depended on physically grounded thought experiments and an internal world model shaped by sensory experience, so text-trained systems should struggle to make the same kind of jump. A follow-up note from the author softened how people were reading it. The paper does not claim AI for science is a dead end, and it explicitly allows that scaling current systems may still produce important discoveries.

Treat this as a useful framing for a gap in current AI systems, not as proof of a hard limit. If you build with LLMs, focus less on whether they can be labeled "creative" and more on what environment, tools, memory, and verification loops let them generate and test genuinely new ideas.

Discussion mood

Mostly skeptical of the paper’s hard conclusion. Readers found the framing thought provoking, but objected that it leans on shaky history, an elite anecdote, and an undefined notion of "jumping" while ignoring how much capability now comes from multimodal systems, tools, and agent scaffolding rather than a bare text model.

Key insights

  1. 01

    Einstein is the wrong clean-room example

    Using Einstein as a pristine example of a lone intuitive leap breaks down once you remember how much of relativity was already on the table. Lorentz transformations, electrodynamics, and Poincaré’s work gave Einstein a large share of the scaffold. What Einstein added was a cleaner physical interpretation and synthesis. That makes the paper’s contrast between human invention and model interpolation look overstated, because the historical case is less "creation from nothing" than "reframing a crowded frontier."

    Be wary of AI arguments that rest on heroic-scientist stories. If you want to test scientific novelty, choose domains where the precursor knowledge and intermediate steps can actually be enumerated.

      Attribution:
    • quantum_mcts #1
    • eru #1
    • kergonath #1
    • kfarr #1
  2. 02

    The claim is not operationalized

    Saying a model cannot make an abductive jump sounds strong, but there is no crisp test behind it. Several commenters pointed out that the paper never defines a measurable threshold for a "jump," never shows what evidence would falsify the claim, and never separates sensory grounding from general reasoning in a way an experiment could probe. Without that, the paper reads more like a philosophical preference than a capability assessment.

    When evaluating AI limitations, ask for a benchmark and a failure criterion. If a claim cannot be turned into a reproducible test, do not use it to guide product or strategy decisions.

      Attribution:
    • killerstorm #1
    • ACCount37 #1
    • nearbuy #1
    • kdavis #1
  3. 03

    The real unit is the whole system

    A recurring useful point was that judging a standalone LLM misses where frontier capability now comes from. Put a model in a loop with simulators, physical experiments, search, memory, and critique, and the relevant question becomes whether the system can propose and kill hypotheses efficiently. That is much closer to how real science works anyway. The missing ingredient may be less "embodiment" in the philosophical sense and more disciplined interaction with an environment that can push back.

    If you want novel outputs, invest in tool use, experiment design, and verification infrastructure. Treat the model as one component in a hypothesis engine, not as the whole engine.

      Attribution:
    • bob1029 #1
    • rsfern #1
    • zamalek #1
    • yogthos #1
  4. 04

    The author’s own clarification narrowed the claim

    The follow-up note changed how many people read the paper. It says this is a personal position paper, not DeepMind policy, and explicitly concedes that current recipes may still deliver important scientific discoveries. That makes the strongest public readings of the title look inflated. The interesting part is not "LLMs are a dead end" but the narrower design question of what architecture would support Einstein-style physical thought experiments.

    Do not let a viral title substitute for the actual scope of a paper. In your own reading list and internal reviews, separate broad marketing interpretations from the narrower technical claim being made.

      Attribution:
    • defgeneric #1
    • wildfireday2 #1
    • throwaway63467 #1
  5. 05

    Language-only training may still miss key world structure

    The most sympathetic thread for the paper was not mystical. It was about representation loss. Several commenters argued that natural language is a compressed and ambiguous record of human experience, while code works better for models partly because semantics are tighter. Even if a model can talk fluently about the Grand Canyon, that does not mean it has the dense multimodal structure needed to simulate being there. The stronger version of this point is not that AI needs human qualia. It is that text alone may be an impoverished substrate for some forms of physical intuition.

    If your use case depends on robust physical reasoning, do not assume more text is enough. Add video, interaction traces, simulation, or instrument data instead of hoping language-only scale will recover everything.

      Attribution:
    • walrus01 #1
    • lukifer #1
    • gabbagool #1
    • air7 #1
  6. 06

    Business value is in augmentation first

    A practical thread cut through the metaphysics. Even if models cannot autonomously produce foundational breakthroughs, they already create value by amplifying human work. Several commenters argued that the industry fixation on replacement is driven more by valuation stories than by product reality. For now the more credible path is systems where humans supply goals, taste, and accountability while models expand search and reduce the cost of trying ideas.

    Plan around human-in-the-loop workflows for longer than AI hype cycles suggest. The near-term opportunity is better leverage per expert, not fully removing experts from high-stakes knowledge work.

      Attribution:
    • aantix #1
    • daun_gee #1
    • goatlover #1
    • juleiie #1

Against the grain

  1. 01

    Recent math results already challenge the thesis

    Some commenters argued the paper aged fast because frontier systems have already been credited with surprising mathematical results, including a unit distance counterexample and work discussed around the Jacobian conjecture. Even if those cases do not settle the philosophy of creativity, they weaken any blanket claim that models are incapable of non-obvious leaps. The burden is now on critics to explain why these do not count without redefining the target after the fact.

    Track concrete research outputs, not just theories of cognition. If AI systems keep producing publishable surprises, your bar for "real novelty" needs to be explicit before the result arrives.

      Attribution:
    • yk #1
    • darkstarsys #1
    • Hammershaft #1
  2. 02

    Next-token mechanics do not rule out creativity

    One sharp rebuttal attacked the common move of inferring capability limits directly from the training objective. Predicting the next token is an implementation detail, not a full account of what behaviors the system can express. The analogy was that comedians also just modulate air with vocal cords, yet that says nothing useful about whether they can be funny. If larger models increasingly show humor, theory of mind, and contextual surprise, then architectural impossibility claims are probably mistaking an abstraction boundary for a hard ceiling.

    Do not overread the base mechanism when forecasting product ceilings. Capability often emerges at the behavioral layer long before theory cleanly explains it.

      Attribution:
    • ACCount37 #1 #2
    • wnmurphy #1
  3. 03

    Human intuition is less magical than advertised

    Several comments pushed back on the romantic idea that scientific leaps come from raw bodily feeling. The equivalence principle can look obvious after you have the framing, or alien before you do. Quantum mechanics routinely defeats everyday intuition altogether. That suggests "intuition" often means trained conceptual habit plus the right representation, not direct sensory grounding. If so, models may need different training, not human-style embodiment.

    Be cautious about product or research plans built on fuzzy claims about uniquely human intuition. Often the bottleneck is representation and practice, which are exactly the kinds of things engineered systems can change.

      Attribution:
    • gus_massa #1
    • CyLith #1
    • kurthr #1
    • dtj1123 #1

In plain english

abduction
A form of reasoning that proposes the most plausible explanation for incomplete information.
Deduction
Reasoning from rules or premises to conclusions that follow logically.
Electrodynamics
The branch of physics that studies electric and magnetic fields and their interactions.
Induction
Reasoning from patterns in examples to broader generalizations.
Jacobian conjecture
A long-standing open problem in mathematics about when certain polynomial mappings have polynomial inverses.
LLM
Large language model, a type of AI system trained on huge amounts of text and code that can generate responses or software from prompts.
Lorentz transformations
Equations that relate measurements of space and time between observers moving at constant speed relative to each other.
Multimodal
Using more than one kind of data, such as text, images, audio, video, or sensor readings.
Special relativity
Einstein’s theory describing how space and time behave for objects moving at constant speeds relative to each other.
World model
An internal representation a system uses to predict how a situation or environment will behave.

Reference links

Author clarification

Physics history references

Reasoning and causality

  • The Ladder of Causation
    Suggested as a better framework than the paper’s argument for discussing what current models lack.

Model capability experiments

Math result examples

Multimodal and chemistry examples

Background explainers

Related cultural references