The paper is a position piece, not an empirical result. It uses Einstein’s path to general relativity as a case study and argues that LLMs are good at induction and increasingly good at deduction, but still lack abduction, meaning the ability to form new explanatory hypotheses from sparse evidence. The core claim is that Einstein’s leap depended on physically grounded thought experiments and an internal world model shaped by sensory experience, so text-trained systems should struggle to make the same kind of jump. A follow-up note from the author softened how people were reading it. The paper does not claim AI for science is a dead end, and it explicitly allows that scaling current systems may still produce important discoveries.
Most of the useful signal was skepticism about how strongly the paper states its conclusion. Several commenters said the historical setup is muddled.
Special relativity did not spring from Einstein confronting a single experimental anomaly. Lorentz, Poincaré, Maxwell, and others had already built much of the mathematical and conceptual scaffolding. That matters because the paper’s chosen example is doing too much work. If the history is messier and more cumulative, then the case for a uniquely human, embodied leap gets weaker.
The bigger complaint was methodological. People objected that "LLMs can’t jump" is presented as a structural impossibility without a measurable definition of what counts as a jump, a benchmark that could falsify the claim, or a comparison class showing that humans reliably do this either. Using Einstein and general relativity also sets the bar absurdly high. Almost no human has ever made that kind of breakthrough. Readers repeatedly landed on a narrower point that feels more defensible. Current LLMs in isolation are not enough. But systems with
multimodal input, tools, simulators, memory, and feedback from an external environment may be. That shifts the question from whether next-token models are philosophically capable of genius to whether AI systems can be built that generate candidate hypotheses and then rigorously test them.
That practical framing also undercut the more theatrical claims in both directions. Some readers pointed to recent math results and chemistry systems as evidence that the line between search, reasoning, and creativity is already blurry. Others said those examples still do not prove genuine novelty. The strongest consensus was simpler than either camp wanted. Today’s models are powerful assistants and increasingly useful components in scientific workflows, but claims of either permanent incapacity or imminent Einstein-level autonomy are both outrunning the evidence.