Rendered at 19:46:24 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
gabbagool 1 days ago [-]
Probably too late for this, but I have argued before that language is a fundamentally lossy encoding of the human experience. We do our best to describe what we're seeing and experiencing using language, which is fantastically expressive, but it has its limits. I think we see glimpses of this when we find ourselves saying things such as, "it's impossible to put it into words" or we overload certain words when we mean very different things, such as, "I love my children" or "I love apple pie". Clearly the word "love" here has a certain magnitude that is not being expressed, yet it is understood by the listener somehow.
So, I do sort of buy into this idea that Einstein was simulating the world and running experiments on those simulations in ways that were beyond what you could encode in natural language. Will AI be capable of doing this, if it is bounded by training data that is composed almost entirely on language? One might argue that if AI is training on a lossy encoding/representation of the human experience, how will it be able to simulate anything beyond that experience? Unless it does so in a way that we manage to do when we image objects beyond 3D. But now I'm just rambling.
lukifer 1 days ago [-]
The lossy compression of language is why we should find it unsurprising that LLMs tend to perform better at code, than at human language tasks or reasoning. While there can be subtle semantic differences in real codebases (using "null" to mean "unknown" in one context, versus "intentionally blank" in another), there is a much tighter coupling of semantics to meaning (low ambiguity) compared to "love" in English (let alone any inexpressible je ne sais quoi).
With apologies if this is common knowledge at this point, 3Blue1Brown has been doing an excellent series on compression, and its relationship to intelligence (or more controversially, that they are one and the same): https://www.youtube.com/watch?v=l6DKRf-fAAM
But that also throws in sharp relief, that there is vastly more to the human experience than intelligence alone: qualia, desire, gut instinct, intuition, emergent creativity. (Whether the "God of the Gaps" for the delta between capabilities of human vs AIs is fixed, or diminishing, or even shrinking to zero, remains an open experiment we're all living through.)
nullbio 10 hours ago [-]
I think LLMs perform better at code than human language tasks because there's no clear way to eval human language tasks in a non-ambiguous or concrete manner. Any sort of eval that happens around language is transformation tasks, which have deterministic properties or goals. The fuzzy side of human communication can only be modeled probabilistically because there are no clear boundaries. Human communication is more like a felt mutual agreement where the correct interpretation is generally determined by popularity.
cwmoore 18 hours ago [-]
I have come across a concept that given a file of compressed text files, adding a new text to it expands it more or less depending on how different the new text is from the compressed.
The compression series sounds interesting!
I think the deltas are growing at different rates.
Hunpeter 14 hours ago [-]
> given a file of compressed text files, adding a new text to it expands it more or less depending on how different the new text is from the compressed.
The 3b1b videos mention how this concept may be used to find similarities between different languages. Researchers have been able to get results that closely resemble how languages are usually grouped into families.
scotty79 7 hours ago [-]
> we should find it unsurprising that LLMs tend to perform better at code, than at human language tasks or reasoning
Math is pure reasoning and stock LLM is already better at it than 99.9% of humans.
And as for human language, agentic LLM can produce a perfectly human text. The fact that AI texts have tells is just because default settings are the same for millions of people using given LLM and almost everybody just pretty much on-shots the text instead of doing it agentically with anti-slop detection and rephrasing in the loop.
walrus01 20 hours ago [-]
> Probably too late for this, but I have argued before that language is a fundamentally lossy encoding of the human experience. We do our best to describe what we're seeing and experiencing using language, which is fantastically expressive, but it has its limits.
Yes, and sometimes this is very intentional. Take for example a short poem which if you sit and really think about it for a long time, you could go off on a mental tangent of imagining what sort of kingdom or empire created a statue that is now "two vast and trunkless legs of stone", for instance. Being terse and allowing for human interpretation is kind of the entire point of something being written like this.
I met a traveller from an antique land
Who said: Two vast and trunkless legs of stone
Stand in the desert. Near them, on the sand,
Half sunk, a shattered visage lies, whose frown,
And wrinkled lip, and sneer of cold command,
Tell that its sculptor well those passions read
Which yet survive, stamped on these lifeless things,
The hand that mocked them and the heart that fed:
And on the pedestal these words appear:
"My name is Ozymandias, king of kings:
Look on my works, ye Mighty, and despair!"
Nothing beside remains. Round the decay
Of that colossal wreck, boundless and bare
The lone and level sands stretch far away.
AnthonyMouse 18 hours ago [-]
> Being terse and allowing for human interpretation is kind of the entire point of something being written like this.
It's also to a certain extent why LLMs work.
When IBM Watson was playing Jeopardy, one of the game prompts was:
> It was the anatomical oddity of U.S. gymnast George Eyser, who won a gold medal on the parallel bars in 1904
The man was missing a leg and used a prosthetic. Watson's output was, "What is leg?"
At first it was regarded as correct. If a human said that you could conclude that they knew the answer. But then the judges decided not to give Watson the point because its output didn't provide enough specificity to prove that it understood the context.
If you ask an LLM what kinds of things taste sweet it can give you examples like cotton candy or strawberries, but it has never actually tasted anything. All it knows is that the training data contains the association between those tokens. But the human reading the output knows what strawberries are, which is what allows the output to be meaningful.
cwmoore 18 hours ago [-]
Human language requires a human receiver, like art requires an audience.
> The fundamental problem of communication is that of reproducing at one point either exactly or approximately a message selected at another point. Frequently the messages have meaning; that is they refer to or are correlated according to some system with certain physical or conceptual entities. These semantic
aspects of communication are irrelevant to the engineering problem.
cwmoore 6 hours ago [-]
Sweet source. I could have said more in one paragraph and 55 pages of formative theory than the two lines above.
cezart 14 hours ago [-]
In this context llm's remind me of the anecdote of Agassiz and the fish:
"A post-graduate student equipped with honours and diplomas went to Agassiz to receive the final and finishing touches. The great man offered him a small fish and told him to describe it. Post-Graduate Student: “That’s only a sun-fish” Agassiz: “I know that. Write a description of it.” After a few minutes the student returned with the description of the Ichthus Heliodiplodokus, or whatever term is used to conceal the common sunfish from vulgar knowledge, family of Heliichterinkus, etc., as found in textbooks of the subject. Agassiz again told the student to describe the fish. The student produced a four-page essay. Agassiz then told him to look at the fish. At the end of the three weeks the fish was in an advanced state of decomposition, but the student knew something about it."
What was the student meant to describe that wasn’t covered in the essay?
morgoths_bane 13 hours ago [-]
No man ever fishes the same fish twice, for he and the fish change. -Heraklitos when eating a tuna sandwich.
andai 19 hours ago [-]
> The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought. The psychical entities which seem to serve as elements in thought are certain signs and more or less clear images which can be “voluntarily” reproduced and combined… The above-mentioned elements are, in my case, of visual and some of muscular type. Conventional words or other signs have to be sought for laboriously only in a secondary stage, when the mentioned associative play is sufficiently established and can be reproduced at will.
—Albert Einstein
Quoted in Using Spaced Repetition Systems to See Through a Piece of Mathematics,
which describes the author's experience that if you approach a field obsessively enough, eventually you begin to understand it at a level deeper than language.
If an LLM is big enough, I imagine something similar is happening.
andai 7 hours ago [-]
Although I imagine that requires contact with actual reality? i.e. I would expect a higher degree of such insights to come from RLVR.
Then again maybe the philosophy department has a different opinion :)
gabbagool 1 days ago [-]
Or, ask yourself the following question:
If you read every single book there is about The Grand Canyon, and watched every single video and/or documentary about The Grand Canyon, do you believe that you have fully experienced The Grand Canyon? Or do you just have to be there to fully experience it.
I dunno. Substitute in whatever you want for "The Grand Canyon". Maybe climbing Mount Everest or walking on the Moon. The point is that maybe the human experience is more vast than what is written about it.
epihelix 15 hours ago [-]
You have to be there to have a subjective experience of the Grand Canyon, for whatever that's worth. (This is essentially the qualia distinction between us and LLMs, writ large.)
But in terms of objective knowledge about the Grand Canyon, reading every article and scholarly work would give you a much better grounding. Unless you're a geologist/ecologist doing actual field research there, actually seeing the Grand Canyon would add little, if anything, to your factual knowledge and understanding of the site.
throwaway0123_5 3 hours ago [-]
Beyond scientific knowledge, being there would (much more quickly) yield quite a lot of (higher quality) knowledge about "how to navigate and do things in the Grand Canyon."
If I'm going to hire a guide, I'm going to hire the guy who spent a week hiking it instead of the guy who spent a week reading about it or watching videos about it.
air7 23 hours ago [-]
In this context that isn't the right question.
Rather one should ask "Would you be able to answer any question and fulfill any request about the Grand Canyon given to you by others in a way that will be indistinguishable from someone who has been to the Grand Canyon"
strken 19 hours ago [-]
This is setting up a contrast between long term memory and an LLM.
You might equally contrast someone who is physically in the Grand Canyon with an LLM given access to a drone with sensory attachments. I'd bet that if you told groups of humans and drone-controlling AIs to find something interesting, the humans would be more variable and cover a wider range of interesting discoveries but often just give up on the task, while the AI would be more likely to find something but would cluster around certain discoveries and make fewer overall.
goodmythical 21 hours ago [-]
What is the current ratio of opened to unopened cans of soda in the canyon, and how has that ratio developed over the last thirty days?
I'd be fantastically surprised if you could provide a precise answer to any such question without extensive security footage that doesn't exist.
An estimate, perhaps, but you probably couldn't even answer a simpler "how many people are currently in the canyon" without resorting to estimation given the lack of information currently available via the mentioned sources.
notduncansmith 20 hours ago [-]
I suspect the same limitations apply to Grand Canyon visitors
Dylan16807 1 days ago [-]
I think I could get very close via video. But that depends on a lot of my non-canyon experience moving around the world, and LLMs are very flawed in their ability to input video, so I think they would not get nearly as close.
throwaway0123_5 1 days ago [-]
I've been to a few natural wonders (including the Grand Canyon) that I saw in advance on video. At least for me it isn't close at all. Even if audiovisual elements could be near-perfectly reproduced by video (imo not even close with modern tech, no screen is capturing the brilliance of sunlight), you aren't capturing the temperature, the feel of wind or rain, the smell of the plants around you, etc.
gabbagool 23 hours ago [-]
Well, I've never been to the Grand Canyon myself, but what you said perfectly aligns with my intuition.
Also, you touched on how even modern tech still falls short of the true experience. Just look at the history of motion picture since the early 1900s. We have added sound, color, bigger screens, faster refresh rates, more pixels, 3D (sort of), surround sound, spatial audio, IMAX, etc. Almost seems like video leaves a lot to be desired.
goodmythical 21 hours ago [-]
4d gaussian splats are going to be a major revolution in vicarious experiences
jhbadger 17 hours ago [-]
I don't know. I think some "wonders" like the Grand Canyon and Niagara Falls are more impressive in aerial videos and are kind of a let down in person -- you can't really see the whole thing when you are next to it.
NegativeLatency 16 hours ago [-]
Mt Rushmore is better and more impressive in media than IRL. Saw North by Northwest a few times and went later in my life just didn’t compare.
Dylan16807 23 hours ago [-]
But you already know those feelings from being outside elsewhere. You get a unique combination there, but that's only worth so much.
throwaway0123_5 22 hours ago [-]
> only worth so much.
Worth... quite a lot imo.
First, in the slightly objective sense of "What is this place like in real life, under X conditions." But, more subjectively, watching a video of a glacier and imagining the wind/rain/cold doesn't even approach 5% of the intensity and awe of climbing up a mountain yourself to see a seemingly endless expanse of ice, struggling to stand steady because of the wind, shivering because of the cold and rain. And finally, in the financial sense, a lot of people routinely spend many thousands of dollars and days-weeks of their time to experience natural wonders in real life.
Dylan16807 22 hours ago [-]
There's so much natural wonder in the world in and around your home.
It's not that the value of the grand canyon is low in an absolute sense, it's that in a relative sense the gap between "empty numb void" and "hiking at home (plus grand canyon videos)" is far far greater than the gap between "hiking at home (plus grand canyon videos)" and "actual grand canyon".
throwaway0123_5 3 hours ago [-]
> There's so much natural wonder in the world in and around your home.
Sure, no doubt. But I don't really think "video of my local state park + imagining the parts that aren't conveyed over video" is particularly close to the real thing either.
AnthonyMouse 18 hours ago [-]
> in a relative sense the gap between "empty numb void" and "hiking at home (plus grand canyon videos)" is far far greater than the gap between "hiking at home (plus grand canyon videos)" and "actual grand canyon".
Have you ever been to a place which is really high up? You're above the trees and can not only see as far as the curvature of the earth allows you to, but are high enough that the curvature of the earth allows you to see for miles. The air is thinner so it's easier to take a breath yet harder to catch it. Your sweat evaporates faster but a breeze doesn't cool you as much. The sun is brighter, rises earlier and sets later. The wildlife makes unfamiliar sounds and sound itself is changed. Even the same food tastes different.
It's not something you can get by sitting at home and changing the background on your computer.
Dylan16807 13 hours ago [-]
I haven't been that high but I've been in places with some pretty impressive views. And the grand canyon doesn't give you that mountaintop experience either.
Back to the topic, I'm sure it's very impressive, but I think you're underestimating just how impressive our baseline is.
jimbokun 23 hours ago [-]
The experience of every famous place I've visited after watching extensive video before hand, has been very different from what I imagined from the video.
pmontra 24 hours ago [-]
No way. You get to the rim and only then you realize how huge it is. And then you still have to go down into the canyon.
dinfinity 22 hours ago [-]
The experience of vastness is mostly due to (high-res) stereoscopic vision. When you experience 6DOF VR in a fairly high-res world it definitely comes much, much closer to the awe one feels with the real experience.
It's similar in dismissing artificial audio based on listening to some pop song on a basic mobile speaker versus listening to well-designed binaural audio on good headphones. The latter can sound 'freakily' 'real'.
arendtio 5 hours ago [-]
I don't think that the 'lossy' part is the problem. Instead, the problem is that language is only a subset of the human experience. So there are aspects that are not included when training models based on language. Making the models multi-modal helps to close that gap, but it still exists.
But ultimately this is just a matter of training data. I do not say that it is easy to obtain the required data, but it is not a fundamental problem LLMs can't overcome.
kkukshtel 16 hours ago [-]
Noah Smith had an interesting idea related to this in a recent newsletter:
> Another way of saying this is that there may be laws of the universe that humans can’t understand but AI can. I call these “cloud laws” — causal regularities that can be exploited by technology, but which are too diffuse and complex for an individual human being to either intuit or communicate.
Human language seems to obey cloud laws, so why not other phenomena too? Perhaps social sciences like economics, sociology, and political science obey similarly complex regularities, and AI can help us find them. Perhaps there are physical processes — plasma, or topological materials, or aerial turbulence, etc. — that obey cloud laws instead of chaos?
I think this is very critical concept and existing gap which does not makes into AI conversations. Interestingly enough a lot of Si Fi movies have captured this where the AI starts to feel and have "thoughts". Have to give kudos to the authors for being so creative and imaginative.
scotty79 7 hours ago [-]
The entirety of our understanding of reality we owe to lossy transformations. Also called modelling.
zapataband1 1 days ago [-]
bro, it is a token predictor. You had me in the first half, language is humanity's greatest invention and all of the LLM "wins" can be attributed to existing language(imo). But no it's not going to have a vision like Einstein because it doesn't have a brain it is a token predictor.
vanviegen 23 hours ago [-]
Yes, and a human is a procreation machine. Bro.
anzuhoeren 1 days ago [-]
[dead]
quantum_mcts 1 days ago [-]
The popular retelling of how Einstein created Special Relativity to "Resolve the contradictions of Michelson-Morly experiments" is very reductive to the history of the question. The epitome is the quote from the paper:
> From the two postulates, Einstein derived the Lorentz trans-
formation ...
If Einstein derived them, who is "Lorentz"?
The groundwork for Special Relativity was the study of electrodynamics and symmetries of Maxwell equations. The Einsteins paper was literally called "On the Electrodynamics of Moving Bodies" and never cites Michelson and Morley.
teleforce 1 days ago [-]
Einstein seems to be so conveniently dismissive that he knew about the seminal Michelson-Morley experiments but he probably knew about it too well [1].
Einstein is not the first great scientist who are in denial of other important prior contributions, and he also not the last one. Newton also probably knew too well about Al-Haytham (Alhazen), arguably the father of modern science, and his breakthrough experiments but never directly cited Alhazen's works in his seminal books on Optics.
[1] Millikan, Einstein, and the Birth of Relativity (4 letters):
"Abraham Pais, who knew Einstein well and wrote his scientific biography, was certain that Einstein did know about Michelson's experiment before 1905. He points out that Einstein was over seventy and in poor health when he spoke to Shankland; at the first interview he probably did not remember that Michelson's experiment is discussed in Lorentz's 1895 monograph, the famous "Versuch", which Einstein had definitely read before 1905."
bryanrasmussen 1 days ago [-]
Was the ethical requirement to discuss and acknowledge predecessors work as developed in Newton's time as in our own? I would naturally expect not because it would seem to me to be the kind of thing that develops over time, but I could be wrong as I am not a historian of science.
seemaze 1 days ago [-]
John Maynard Keynes is famously quoted as stating:
“[Newton] was not the first of the age of reason. He was the last of the magicians.”
notpachet 1 days ago [-]
We seem to have come around to a new magical age.
teleforce 1 days ago [-]
It's a not a requirement, it's just a natural to do as a research scientist and a student of knowledge to acknowledge prior contributions if you knew it. The same is true when you are at the time of Aristotle.
bryanrasmussen 22 hours ago [-]
while I said I was not a scientific historian I am well enough read that I can say no, the tradition of citing all sources is a relatively modern invention. Definitely not widespread in Aristotle's time, and not extensively practiced by Aristotle himself who very seldom names any source and their work, but I was not sure if in Newton's time it was already beginning to be established as I have not actually read many scientific papers from that time (some philosophical, but not science)
Furthermore I can't help but feel some logic would lead us to conclude it is not natural to cite, why? Because in our time if you do not cite your sources in a lot of occupations dealing with knowledge you will be punished for it. It would be absurd to construct punishments for failing to do what people will naturally do, and to have those punishments so often called on.
Even if we give a very lenient interpretation of not citing sources we would probably have around 10% of papers with errors, and that is with the fear of punishment.
Citing when one should and accurately is not normal, it requires discipline to do it, or there would not be such a high failure rate at doing it.
schubidubiduba 22 hours ago [-]
You need to differentiate between citing everything, and citing a select few prior works which are integral to your own. For the latter, it is very natural to mention and cite them, unless you have motive not to do so.
bryanrasmussen 13 hours ago [-]
that, however, was not what was said - people were arguing about people not citing, and Newton not citing, and I asked if it was expected to cite them at that time. At which point another person argued that yes it was and that even back in Aristotle's time it was natural to cite, which the use of Aristotle, the guy whose predecessors seemed to almost all be either "some guy", "people say", or "those pesky Pythagoreans", as an example of someone who was naturally prone to cite put me a bit on edge.
So it seems you have established that no, Newton was not expected to cite everyone in his time, very well, this now puts the question further up the chain over my entry into the conversation - "Were the people Newton neglected to cite actually integral to his work?"
My definition of integral would be, two possible definitions:
1. do I write something here that someone reading it would say, "What?! What do you mean, explain yourself, how did you arrive at this surprising bit of information?!!" then it is integral.
2. Am I doing this based on some work that someone did without which I would definitely not be doing this? If I am doing it with knowledge of the prior work but whether or not the prior work existed I would still be doing it, the prior work is not actually integral.
randallsquared 1 days ago [-]
"Forgive him, for he believes that the customs of his tribe are the laws of nature!"
WorldMaker 24 hours ago [-]
It reminds me a bit of an anthropology question I've heard brought up a lot that's sort of the dark inverse of xkcd 1053 [0], that many things that "everybody knows" tend to get lost over enough time. You don't need to properly cite experiments or lectures or books that "everybody knows". Especially when the work you do essentially replaces those things in the canon, such connections sometimes get entirely lost. To an anthropologist working in some ancient language the variations of the phrase "as everybody knows" are often a missing or lost citation today. The places where it is so implied that you have the same education or background or textbook as the author to not even warrant an "as everybody knows" is an ever worse thing because you don't even have the marker that you are missing a citation.
See also this article, which discusses transcripts from lectures given by Einstein in Kyoto and Chicago in 1922 and 1921 (respectively), and letters from 1899, with earlier accounts of the Michelson-Morley experiment influence:
This isn't offered as a defense of Einstein since I couldn't speculate on what ideas he must have been consciously aware of, but several times in my life I've thought I had a novel idea only to find a very similar idea was in something I read and had just forgotten about. It's a very human phenomenon.
Science rightfully recognizes the mind that doesn't just first describe the idea but provides a robust framework to test it and communicate it.
eru 1 days ago [-]
I agree with you in general. However:
> If Einstein derived them, who is "Lorentz"?
You can (re-) derive a lot of existing stuff.
kergonath 1 days ago [-]
> You can (re-) derive a lot of existing stuff.
Einstein was aware of Lorentz and the transform. He was aware of Poincaré as well. He knew the state of the art for his time.
eru 1 days ago [-]
Re-deriving the start of the art cleanly with fewer postulates is good!
oliculipolicula 1 days ago [-]
And so I'm skeptical of TFA's claim that "creativity _isn't [just]_ compression"
Imho there's also a clean argument against the existence of LLM-understanding: chatbots have been unable to summarize to experts (see my reply to you in the other thread) their own findings.
Even after prolonged interrogation. They were unable to _compress_ their own findings. Thus they might not actually understand what they have actually done. (They might barely pass an oral thesis defense)
(One may object-- that proofs aren't data that can be "compressed". But then doesn't the process of abduction generalise the very idea of data? to.. ?)
eru 1 days ago [-]
I'm not sure, many chat bots summarise just fine as far as I can tell, on subjects I know a lot about.
1 days ago [-]
jimbo808 1 days ago [-]
I get really tired of Einstein being portrayed as if he’s the greatest scientist who ever lived, when he doesn’t even belong in that conversation. If Einstein never existed, the fields he was in were already moving strongly in the directions of his conclusions. If Maxwell or Newton had never existed, the world today would be quite different.
dguest 1 days ago [-]
Yeah, once you dig into economics, technology and science of the era it often feels like the big discoveries were inevitable. That doesn't make these people any less great, it's just to say there are many more great thinkers that people are forgetting about.
But I'd like more explanation as to why Maxwell or Newton are any less replaceable: all 4 of "Maxwell's equations" have other names attached to them (the exception is Ampere-Maxwell law [1] which Maxwell contributed an important term to), and Newton had a number of contemporaries who were making similar discoveries but are often forgotten.
Einstein did a lot, he not only published the special and general theories of relativity, he also contributed to quantum physics and was the first to deduce the size of an atom using statistical arguments, essentially proving the existence of atoms.
adrian_b 1 days ago [-]
Not at all.
The size of an atom was computed for the first time 40 years earlier, in 1865, by Johann Josef Loschmidt, who was an Austrian scientist.
Loschmidt had determined the value of what today is called "Avogadro's number", despite the fact that Avogadro had no idea about how one could find the value at that number. The contribution of Avogadro had been the law that ideal gases have the same number of molecules per volume when the pressure and the temperature are the same, but he did not know how that number might be determined.
Based on the results of Loschmidt, which were also used by Maxwell in his works on statistical physics, a few years later, in 1874, George Johnstone Stoney computed the value of the elementary electric charge, which he later named "electron".
Loschmidt was the one who first measured and weighed the atoms, establishing their existence beyond any reasonable doubt. Before him, many believed that the hypothesis of Dalton about the existence of atoms is just a convenient theory for explaining the laws of chemistry, but some other better explanation might be discovered in the future.
Einstein published an explanation of the Brownian motion, which verified that everything fits in the known theories and it matches the expected values. It did not contain any novel number or law, but it was a useful confirmation of the current theories.
In my opinion, the greatest achievement of Einstein has been the paper published by him in which he revealed the existence of the stimulated emission of electromagnetic waves by hot bodies, besides their well known absorption and spontaneous emission.
This paper had tremendous practical applications, leading to the discovery of masers and lasers, without which the modern technology would have never existed.
While Einstein's theory of gravitation was more intellectually challenging, its practical applications are minor, almost negligible. Moreover, even after more than a century we do not know exactly how adequate it is for describing gravitation.
What is certain is that Einstein's theory of gravitation cannot be anything else but an approximation, as it is based on averaging the distribution of matter in space, without taking into account the concentration of matter into elementary particles.
Another contribution of Einstein that is more important than his most frequently cited works is his support for Satyendra Nath Bose, which lead to the recognition of the Bose-Einstein statistics, which together with the Fermi-Dirac statistics plays an essential role in quantum physics.
naasking 6 hours ago [-]
I don't understand all this dismissal of Einstein. What you just described is that he published a unified theory of Brownian motion, developed a new way of conceptualizing scientific theories via principles with which he derived special and general relativity, contributed foundational work to quantum mechanics, and more, but no, he was not one of humanities' greatest scientists.
adrian_b 2 hours ago [-]
Have you even read what I have written?
What I have written is not a dismissal of Einstein. On the contrary, I have said that Einstein is the source of a few extremely important discoveries, but those are not the theories cited by most people, who have never read the original work of Einstein but speak about him from hearsay.
As I have written above, when speaking about Einstein, almost nobody mentions the discovery of stimulated emission, which has been far more important than anything else done by Einstein. Einstein did not discover special relativity, the photoelectric effect or Brownian motion. He just provided alternative explanations for them, which did not include any new formula or quantity.
On the other hand, stimulated emission was a surprise, even if it is one of those things that appears to be obvious in retrospect, but for decades nobody who studied the blackbody radiation had noticed that something is missing in the equations.
This discovery eventually lead to the invention of masers and lasers and today there exists almost nothing that incorporates modern technology and which could exist without lasers.
Already for decades, no modern integrated circuits, no CPU and no memory can be fabricated without using lasers. Even many purely mechanical things, but which require great precision, cannot be fabricated without using measurement instruments that depend on lasers.
So the contributions of Einstein are very important, but they are not those that most people associate with him. Besides stimulated emission his most original work was his theory of gravitation, a.k.a. general relativity, but if he had not developed it that might have changed nothing, because Hilbert had developed an equivalent theory almost simultaneously, so had Einstein not published first, Hilbert would have become its discoverer.
After stimulated emission, the most important in practice research of Einstein is the Bose-Einstein statistics. But again, here like for special relativity, the statistics had been discovered by someone else, i.e. by Bose. Einstein has nonetheless the great merit of understanding its value. He demonstrated that the statistics was much more generally applicable than initially believed by Bose and his authority convinced everybody about its importance.
Thus there is no doubt that Einstein is one of the most important physicists of the 20th century. Nevertheless, there are dozens who made similarly important discoveries, e.g. Sommerfeld, Dirac, de Broglie, Schroedinger, Fermi, Rutherford and many others.
potamic 6 hours ago [-]
Leading physicists at one time appeared to disagree. Newton and Maxwell do come at #2 and #3 though.
That voting has obviously been done by people who have never read the original texts of the people for whom they have voted.
Einstein is indeed one of the most important physicists of the 20th century, but it is really meaningless to try to establish a hierarchy between physicists, because their works are like open-source programs, each work is based on the works of the predecessors and it would have been impossible to do without those previous works.
In any case, Einstein was good, but Maxwell seems really superior in the sense that he could extract the essential information from the much more confuse sources that were available at his time.
The Treatise of Electricity of Magnetism of Maxwell and his collected papers are very instructive lectures even now and there are fragments of them that remain better than the equivalent fragments in many modern textbooks, despite the huge additional information that we have today.
Most of the original texts written by Einstein are now of interest only for understanding the details of how his thinking and that of the contemporaneous physicists have evolved during those important years, but those texts have little other importance except for history, because later better expositions of the same material have appeared. Nonetheless, it is important to also read Einstein's originals, because in many later books there are some inappropriate interpretations attributed to Einstein, but in reality Einstein had said something else and which was more correct than what some later author misunderstood.
john_strinlai 1 days ago [-]
>I get really tired of Einstein being portrayed as if he’s the greatest scientist who ever lived
fair, there's lots of really awesome scientists, many of which not talked about even 1/100th as often as einstein
>when he doesn’t even belong in that conversation.
suddenly, the pendulum has swung way too far in the other direction.
kskdkwkdjw 1 days ago [-]
Oh boy. I just cannot believe we have reached a point where someone would find it acceptable to have a temper tantrum over Einstein’s fame.
What a depressingly peculiar time to be alive.
Reddit is over that way, my friend. You might find the crowd there more amenable to this nonsense.
jimbo808 1 days ago [-]
Calling my comment a “temper tantrum” is the exact kind of behavior you’d expect of a Redditor. You may be projecting.
Einstein’s fame isn’t the problem. It’s the narrative that he is somehow the greatest scientist ever, when it’s just transparently false when you look at his actual contributions and the historical momentum of the fields he contributed to. There’s nothing wrong with his fame, the problem is that it overshadows actual juggernauts, like Maxwell and Newton. Maxwell not being famous at all among the ordinary public is a great tragedy, when he genuinely is in the conversation as having have been the most important physicist to have ever lived.
defgeneric 1 days ago [-]
Worth reposting a follow-up tweet from the author Tom Zahavy [1] after this made the rounds on X/Twitter recently:
> A few reflections on my "LLMs Can’t Jump" paper:
> My position paper recently got some traction here, so I wanted to share a few thoughts and clarify a few things.
> First things first: some people are framing this as "DeepMind is throwing cold water on AI for science" or claiming the paper argues LLMs can never make real scientific discoveries. This is NOT the case.
> This is a personal position paper, not the company's view on AI for science. This is also not my position. As a core contributor to AlphaProof (the first AI system to win an IMO medal), I know firsthand that my colleagues at DeepMind, other frontier labs, and academia have made amazing discoveries with LLMs and will continue to do so. This paper is NOT an "LLMs are a dead end" kind of thing.
> Rather, the paper is the result of a deep dive I took to study the invention of General Relativity. I wanted to explore what it would take for a modern AI system to make that exact kind of jump. Specifically, I focused on the equivalence principle—a key axiom that Einstein formulated through thought experiments grounded in his physical intuition. I was trying to figure out what it would take to give modern AI systems that sort of thinking.
> Giving AI this specific capability isn't necessarily the most urgent thing to do next. It is very likely that improving our current recipes will lead to many exciting discoveries in the near future. In fact, that is what I am personally working on these days (sorry to disappoint you!). It is also quite possible that I am wrong, and that simply scaling our current systems will lead to new inventions in physics and elsewhere.
> Nevertheless, this was my position last winter when I wrote the paper, and I'm sticking to it. I think that there are a few interesting ideas to explore in this space which could influence the next generation of AI systems. I was very lucky to receive a lot of interesting feedback about this position—thank you for all the messages!
>> Specifically, I focused on the equivalence principle—a key axiom that Einstein formulated through thought experiments grounded in his physical intuition.
It's weird because the equivalence principle is very unintuitive. Aristotle's Mechanics does not have it. It took almost two thousand years to discover inertia that is the most simple version of the equivalence principle. Einstein understood the idea of the the equivalence principle because he had a physics degree, not because he feel that in real life.
Moreover, if you ever have to study or teach Quantum Mechanics, physical intuition gets in the way. A lot of properties contradict the physical intuition but after a while you get use to them. If we continue with Einstein, the photoelectric effect does not aperar in real life.
glitchc 1 days ago [-]
> If we continue with Einstein, the photoelectric effect does not aperar in real life.
Of course it does. How do you think your phone camera works?
gus_massa 23 hours ago [-]
But he didn't have one :) I guess unless you use the raw images, I think the cameras fix a lot of the noise caused by the Poisson distribution. I imagine a movie where Einstein get bored in the patent office and decide to open a photography shop, perhaps with very grainy underexposed photography is possible to notice the quantization of photons, and then he gets a Nobel prize.
Another bad idea for a movie is a watermillpunk universe, where during a practice on a hot day the best ever curling player discover inercia.
glitchc 5 hours ago [-]
There were already prior reported cases of the phenomenon. Einstein provided the mathematical foundation explaining it, in particular that the discharge is quantized, indicating that it came from an electron. Millikan proved it a couple of years later with his oil drop experiment and by measuring the elementary charge of an electron, garnered his own Nobel:
That’s funny, when I first learned about the equivalence principle, my first thought was “of course!” I have always found it to be very intuitive. The great leap is being able to frame it that way.
kurthr 1 days ago [-]
I worked on the early iPod scroll wheel, before there were advertisements for it and it was in common use. I found the UI interaction odd and unintuitive, particularly the "menu" button and the dead end you hit when "playing". Of course by the time the demonstrator ads came out and everyone was talking about how easy it was use, I'd already spent 10s of hours on it, and it WAS second nature. Being first and embedded in the culture of the time has huge UI advantages.
(Note that the HP Chipmunk 9836 also had a scroll selector wheel in 84)
Most people are UI bigots. Once they get used to a first something, they expect everything to work that way and hate learning anew. They get stuck on keyboards, mice, trackpoint nubs, trackpads, trackballs, scroll wheels, or touchscreens and refuse to move on. Of course there are 'objective' performance tests for each including Fitt's test of accuracy and latency, as well as, cognitive load. I guess once you have a hammer, every screw looks like a nail.
So I'd be a bit careful with the "of course!". It may well be obvious only because that is the first mental model you latch onto.
wildfireday2 1 days ago [-]
Moreover anyone glomming onto this paper for goal-post-shifting “AI can never” should:
The worst thing about the AI boom is how tech bros feel comfortable abusing the goalpost fallacy. So annoying.
wildfireday2 1 days ago [-]
If your glib comment is referring to me as a techbro and doing the annoying worst thing, then maybe you should explain how I am comfortably abusing the goalpost fallacy given the explosion of agentic AI capability.
I simply point out here, that fully accepting the paper’s premise, the paper’s conclusion isn’t limiting on frontier AI reasoning agents. The paper posits the necessity of multimodal world models and the limitations of LLMs. Frontier agents aren’t simply LLMs and do increasingly integrate increasingly capable multimodal models.
> Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
> Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.
> When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
throwaway63467 1 days ago [-]
I mean Einstein had help, he was networked with the best scientific minds of the planet and his discoveries were grounded in experimental results that contradicted existing theories (at least for specialized relativity), and without Riemann’s work he wouldn’t have been able to formulate his theory either. So not sure if AI couldn’t do that if you kept feeding it with new research results and let it correspond with top human scientists. Einstein was a genius but I don’t think his thought process is beyond what an LLM could simulate. And again this is probably the most impressive scientific achievement in theoretical physics in the 20. century so maybe it’s hanging the bar a bit high for LLMs.
jvanderbot 1 days ago [-]
Came for:
"A computer once beat me at chess, but it was no match for me at kick boxing."
TFA was actually about leaps of intuition, sadly.
One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.
ben_w 1 days ago [-]
> One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.
Could be, but preventing leakage from more modern stuff can be challenging.
I can't find the citation right now, but I think people found it was leaking anachronisms? So this probably wasn't as well filtered as the creator had hoped?
morkalork 1 days ago [-]
Progress followed improvements in hardware, would you have to give access to modern hardware in the experiment for it to use? How much could it infer from it?
ben_w 1 days ago [-]
> would you have to give access to modern hardware in the experiment for it to use?
At a minimum, yes. IIRC, the sum total of all compute manufactured over history only reached the minimum needed to train an OK LMM in the mid 00s.
> How much could it infer from it?
Only way to find out is to try.
handoflixue 15 hours ago [-]
They did that with "Talkie", a model trained on 1930 and before. It had the ability to assemble crude Python programs, but there were definitely some leaks in the training data (it knew about stuff like World War 2) so not 100% perfect. Still seems like a pretty reasonable "proof of concept"
> One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.
This is an interesting experiment but I wonder if it would be possible to prevent some sort of retrospective bias. For example, I’d expect the experiments that lead to relativity to be over-represented in our catalogue of scientific literature prior to 1900, just because in retrospect they were important, so the records about them were preserved.It would have to be a very intentionally constructed corpus, I think.
elar_verole 1 days ago [-]
I think this can't work because an LLM needs too much data, and before the internet there probably just wasn't enough to get close to what we have now
jvanderbot 1 days ago [-]
Even simpler: Can GPT-2 anticipate and build Gwen/Deepseek? I think the answer is almost trivially "no", so I wonder what changed?
ben_w 1 days ago [-]
Lots of things changed, GPT-2 is small (1.5e9) and is also a base model, so it is only doing next-token/autocomplete rather than prompt-response like even the first ChatGPT-3.5 was doing.
jebarker 1 days ago [-]
Just for the sake of clarity: all LLMs up to today are still only doing next-token/autocomplete. The training process got additional stages to shape the model weights, but standalone LLMs are still deployed essentially identically.
ben_w 1 days ago [-]
If you gave GPT-2 a question and ended with a "?", it might answer, but also it might write several more questions in a similar category.
IMO, the mechanism isn't the important thing, the behaviour is. If you look at the step-by-step, we are also looking for the next word or motor action (and for whoever is about to suggest that we humans plan ahead, Transformer-based LLMs have been shown to also do this); as this is not a useful description of what it means to be a living brain, I'd say it's also not a useful description of what makes everything post-InstructGPT different from what came before.
jebarker 1 days ago [-]
I agree completely - behaviorally the models have changed drastically due to RLHF, RLVR and now maybe even more so due to agentic harnesses. But the mechanism of prediction hasn’t changed, that was all I was clarifying.
joefourier 1 days ago [-]
What about multi-token prediction and speculative diffusion? That’s a different mechanism of prediction, even if it serves only to accelerate decoding.
jebarker 1 days ago [-]
As you say, that's just an efficiency play and, as I understand it, doesn't change the behavior of the models beyond perhaps a small amount of sampling noise.
wizzwizz4 1 days ago [-]
If you frame it like so:
<noob> Where do birds go when it rains?
<expert> They
then GPT-2 generally doesn't write more questions.
ben_w 1 days ago [-]
Generally. Sometimes it still did, in my experience.
rowyourboat 1 days ago [-]
That's not really a fair comparison, no? Modern LLMs are much more capable than GPT-2. We'ld need a modern LLM trained on exclusively old data, and that might be impossible
inigyou 1 days ago [-]
Why couldn't an LLM, if it was smart enough, generate and consume its own data?
I know the answer: because it leads to model collapse. But why is that? Wouldn't a smart model not collapse? It's seeming like they keep getting smarter because we keep pouring more of our own knowledge into them, not because they are actually getting smarter. And yes, sometimes a dumb but persistent bruteforcer can make new discoveries.
Garlef 1 days ago [-]
> if it was smart enough
and i think this is exactly the crux;
the really big models need really big datasets
and current gen LLMs get a lot of training data beyond "all books + all of the internet"
the objection is then that producing this additional data would already confound it with pre "virtual cutoff date" knowledge (since the training data probably implies mathematical and SWE concepts that were developed post "virtual cutoff date")
Plasmoid 1 days ago [-]
It's because LLMs are entropy generators. That's not a bad thing for what people are doing.
But to prevent model collapse you need a way to pump down the entropy. Much like in thermo, it's an expensive and slow process.
smusamashah 1 days ago [-]
If it is smart enough to generate data it can consume to train itself better, it is already smart enough to not need to do that.
inigyou 1 days ago [-]
If a human is smart enough to do the Michelson-Morley experiment, they are smart enough to not need to do that.
Kinrany 1 days ago [-]
They are already trained on generated data I believe
hackernudes 1 days ago [-]
Maybe we can synthesize large amounts of limited information. I thought that new training data is mostly synthetic anyway.
naasking 4 hours ago [-]
Typical LLM pretraining is very inefficient with data. NanoGPT slowrun shows that data can be used much more efficiently.
Normal_gaussian 1 days ago [-]
The curious case here is how much of a description do we give it of itself? That would almost certainly dominate success rates.
My feeling is that a prompt would have to provide a vague description of a program that meaningfully passes something like a Turing test, an API to conform to, an expectation of novel construction (no 'ifs all the way down'), and then a requirement to search broadly and pursue promising ideas and not get hung up on the philosophy. Anything more precise feels like it would corrupt the test, but as it is that description feels doomed to loop before even trying the interesting parts.
ModernMech 1 days ago [-]
I wonder if we could just tell it to invent itself without any description and see if it can I introspect enough through its own interface to figure out what it is.
unfitted2545 1 days ago [-]
If would be interesting to see 5 billion LLM's working together, each with random mutations (temperature ig). Would we essentially be looking at a society through a petri dish? Ofc 5 billion is quite a lot of compute.
Tade0 1 days ago [-]
Is that how chessboxing was invented? Genuinely asking.
the_af 1 days ago [-]
Nope.
Chessboxing was the invention of comics book artist Enki Bilal (and he's credited with this in Wikipedia). I first saw it in his Nikopol trilogy. Because life is weird, it then became a real thing.
It's unrelated to computers playing chess. It predates Kasparov's first defeat by Deep Blue. I don't remember any mention of computers being good at chess in the trilogy, either. Or any computers, for that matter.
ambicapter 1 days ago [-]
Funny, I knew about chessboxing and Enki Bilal, but had no idea one birthed the other.
throwaway314155 1 days ago [-]
Article was plenty interesting to me.
killerstorm 1 days ago [-]
This is literally an opinion of one dude which is not backed by any kind of quantitative evidence.
It's actually possible to answer this question rigorously:
1. Define a scientific result which qualifies as a "jump". They should be frequent enough that they happen every year - otherwise one might say humans can't jump either.
2. Identify all such "jumps" in articles published in 2026, and use LLM with 2025 knowledge cut-off to re-derive these results with minimal amount of information.
It really irks me that people boost these low-effort articles just because they confirm pre-conceived notion that LLMs are limited
NewEntryHN 1 days ago [-]
You don't need any reproduction. Assuming researchers are using LLMs, you should just see the number of jumps increase as models get better.
andai 1 days ago [-]
I boost it so someone else will see it and prove it wrong.
emp17344 1 days ago [-]
LLMs most definitely are limited. What’s your position, that LLMs have no limits? That’s obviously wrong, and the fact you hold such an unreasonable position may be why you react so strongly to pieces like this one.
killerstorm 22 hours ago [-]
What's the obvious limitation?
To Yann LeCun it was "obvious" that you can't build a world model from a text, so he predicted that even "GPT 5000" can't predict that an object sitting on the table would move together with the table. But GPT-4 could solve this task and provided correct explanation. GPT-5 can solve much more complex tasks, code physics simulation, etc.
Now people talk about invention of General Relativity as something LLMs can't do... something almost no human was able to do either.
JacobAsmuth 20 hours ago [-]
You're hallucinating facts again.
sobiolite 1 days ago [-]
The theory is that creative leaps in theoretical physics require a grounding in sensory experience, but the obvious counter-argument is that humans can make creative leaps in abstract fields without such sensory grounding. They do address this at the end, saying
"In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such as generality or minimality."
But if such sense experience is possible in abstract domains via some high-dimensional topology, why could a sufficiently advanced LLM not develop an equivalent high-dimensional topology for domains like physics and use it to make creative leaps?
nearbuy 1 days ago [-]
The paper also fails to show that their central example, Einstein, relied on sensory experience for his intuition leaps rather than general reasoning. They just kind of claim that thought experiments require sensory experience. But you can easily ask an LLM to perform a thought experiment and simulate an outcome, and the SOTA LLMs generally seem to do about as well as a human. An LLM would certainly know that freefall feels the same as zero gravity, even if they haven't felt the sensation, which was the key intuition the paper talks about for Einstein's General Relativity. The paper's author would probably say any examples of this don't count, but without clear criteria for what would count, their claim is unfalsifiable.
bob001 1 days ago [-]
That's an interesting analogy. My gut sense is that theoretical mathematics requires a high level of intelligence versus more grounded domains. That may imply that deficiency in grounding can be made up for with intelligence and basically reverse engineering the gaps in grounding from first principles/limited grounding. The ultimate question would then be what is the tradeoffs between grounding and raw intelligence for the same outcome.
Spacecosmonaut 1 days ago [-]
Isnt the sensory grounding even in abstract cases some (limited) intuition that simulates in a mental world model?
roenxi 1 days ago [-]
> ...but the obvious counter-argument is that humans can make creative leaps in abstract fields without such sensory grounding...
But we have no idea at how good humans are at that. Given the appalling failures of humans to handle even basic statistical situations like identifying that the same thing happens over and over, it might be that they are hilariously bad at creative leaps in abstract fields, it is just we have had nothing better available to measure against. We've spent about as long as decision theory existed trying to convince people to use it instead of flailing. Limited success, usually in exceptional cases.
And the paper seems a bit dodgy, we have models created with sensory data available. No reason a LLM can't be trained on more sensory data than a human can accumulate in one lifetime. There is a lot of visual data on YouTube.
wongarsu 1 days ago [-]
Most humans are probably bad at it. Some humans are very good at it. I'd even argue most science does not demand these kinds of leaps and is mostly concerned with incremental improvements, or proving or disproving other people's abductions
The case study they chose is literally Albert Einstein coming up with General Relativity, something most scientists of his time were not able to do
SkyBelow 1 days ago [-]
If there is enough cross over between real world knowledge engrams and abstract knowledge engram, would this allow for the jump?
One interesting (albeit sad) area which might be related are humans who are never raised with a first language. They seem to never developer abstract reasoning and even seem to lose the ability to develop it later in life. This might indicate there is some 'real world senses' -> 'direct language' -> 'indirect language' -> 'abstract abduction' hierarchy that develops, perhaps related to more real world abductions as a necessary side chain to developing abstract ones.
One of the obvious problems with this is just how difficult we find it to study intelligence purely in humans. We are measure a LLMs by a yardstick that is already known broken, but maybe this is still the right path.
yomismoaqui 1 days ago [-]
Why everybody is obsessed with replacing humans with LLMs when it seems like the most profitable use cases (like coding agents) rely on enhancing human capabilities?
Until LLMs have some 0% error humans will have to be in the loop (even if they only serve to take responsibility of the process).
daun_gee 1 days ago [-]
[dead]
StevenWaterman 1 days ago [-]
> The only way to justify trillion dollar valuations
Also possible if you make god
daun_gee 1 days ago [-]
[dead]
goatlover 1 days ago [-]
Would have been better if AI was known as Augmented Intelligence as it's a tool to to boost our own intelligence. Instead the term that promises science fiction futures predominated, and now it's being used to raise a ton of money.
Saying you want to make workers more productive and provide better tools for people just isn't that sexy.
Mikhail_Edoshin 16 hours ago [-]
True understanding requires not knowing, and LLMs cannot "not know". LLMs have to come up with an answer, this is their nature. They are search engines. We do have a similar mechanism; one can notice it by reflecting. The mechanism is an opposite of true thinking, as it merely looks up what is already "known". We "jump" when we temporarily turn this mechanism off.
That said, here's an experiment conducted by some Soviet psychologist, I forgot the name. The man wanted to study intuition. So he invented an experiment that was supposed to trigger it in laboratory conditions. (Take a moment to marvel at that; how would you approach such a task?) He gave people a few puzzles. One was to place some sticks according to some rules. Yet another was to find a path in a maze. The secret was that the path in the maze was the same figure as the solution to the stick puzzle.
And he observed interesting results. People who solved the maze after the sticks found the path much faster than the control group. If a subject was asked to comment how he was solving the maze, at the start or halfway through, the speed dropped to typical. Subjects normally didn't notice the similarities.
So there is something to study here, although it is obviously a case of pattern matching, only subconscious. This is a jump of sorts, but not the one I mean. What I mean is a Zen jump.
yk 1 days ago [-]
The paper is from the 30th of April this year, openAi announced the counter example to the unit distance problem on the 20th of May. That is to say this paper seems to have aged not much but quite poorly.
Hammershaft 20 hours ago [-]
I'm skeptical a counterexample is good evidence of creative intuition.
bob1029 1 days ago [-]
I think this is more of a function of the harness and the environment than the LLM. I've seen some LLM interactions over complex environments like Godot and Unity that challenges the notion that there is no "jumping" going on at all.
An LLM in isolation from its environment might as well be a brain in a vat in some dark cave. You need an external environment to sample from and act upon to make forward progress.
doginasuit 1 days ago [-]
I'm also amazed at the degree an LLM can get the drift with a vague or incomplete prompt. The ability to perceive and operate based on patterns that go beyond the language in the text makes it seem like they would be unusually good at taking leaps that haven't occurred to us.
altmanaltman 1 days ago [-]
Why does the cave need to be dark if its just a brain in a vat?
Maybe LLMs can't, but another form of AI will. I hope nobody is interpreting this as "nothing will never be as good as us".
I see similar thinking in stories of how humanity got here. Religion has thousands of years adapting to this problem, every time we explain something, the goal post moves. Catholics today accept evolution (or least the church does), but it is the "jump" from monkeys to humans where God is the only explanation.
Just 5 years ago we didn't have a technology that knows more about everything than even most experts. We keep coming up with benchmark after benchmark and LLM/AI keeps destroying them. Now we've moved the benchmark to "the jump". Again, maybe it's LLMs or the way we currently do them that can't do this, but eventually something will.
bigstrat2003 1 days ago [-]
> Just 5 years ago we didn't have a technology that knows more about everything than even most experts.
We still don't have that. LLMs have shown time and time again that they don't know a single thing and are incapable of reasoning.
naasking 4 hours ago [-]
No such thing has been shown.
kskdkwkdjw 1 days ago [-]
> Just 5 years ago we didn't have a technology that knows more about everything than even most experts.
And we still don’t. What we have are simply very advanced search results aggregators with delusions of personality. Just because your fridge says “I” doesn’t mean it is a person.
The article seems like an interesting Gedankenexperiment. However, I think it overrotates on the GR analogy.
For example "..ARC captures the logical leap, it misses the manipulative component—the physical sensation and embodied simulation..." makes lots of assumptions on how such a discovery must occur, e.g. through "physical sensation and embodied simulation". Results matter, not the path there.
For example, quantization of energy, at the core of QM, wasn't discovered through "physical sensation and embodied simulation" at all. Planck simply found that if energy is quantized, then one obtained the observed black-body radiation spectrum. There was no "physical sensation and embodied simulation".
wnmurphy 1 days ago [-]
I had a related insight, but in the domain of humor [1].
LLMs are inherently probabilistic, and there's currently no mechanism for producing an orthogonal directional change in the path traced through a latent space which is also contextually relevant (landing on a punch line).
In other words, LLMs are fundamentally incapable of making intuitive/orthogonal leaps in context.
It might be possible to add this capability with a new architectural component like transformers, but specifically for making "left turns"/intuitive leaps.
You have committed the classic blunder of confusing your abstraction layers.
"Probabilistic next word prediction" and "humor" sit about as far apart as "modulating airflow with meat flaps" and "humor" do. One is an interface through which an action is performed and the other is a highly abstract capability.
Would you claim that a podcast comedian is fundamentally incapable of being funny because all he ever does is wiggle the air with his throat meat flaps? Probably not.
Absolutely nothing about "probabilistic next word prediction" forbids "making intuitive/orthogonal leaps in context". The interface is expressive enough.
And empirically? The "sense of humor" in LLMs is yet another "a function of model scale" capability. GPT-4.5 was reportedly funnier than both GPT-4o and o1. Fable 5 is reportedly funnier than Opus 4.x. It's one of those ever-elusive "big model smell" signs that are hard to measure with anything other than vibes.
Under the "humor as an opposed social intelligence test" family of hypothesis, what "being funny" reflects is the funny guy's ability to model and predict you and your reactions. For the comedian to be able to make the audience laugh, he must know his audience well, model it accurately enough to be able to spot the "breaking points" of humor, things they'd find unexpected and clever and thus "funny", and then weave those things into the jokes.
Then, a bigger LLM gets better at humor because it has a more accurate model of how humans think of things - including the "ha-ha" gaps. It's a "theory of mind" capability. It's not "special", it's just hard.
wnmurphy 1 days ago [-]
Humor requires a sudden orthogonal leap from context. That's what a punchline is.
I think you're saying that you can eventually train models to arrive at that destination by training on existing jokes, effectively encoding these leaps as probabilities.
In that case, the model isn't actually making an intuitive/comedic leap; they're just following new probability chains in attempting to approximate examples they've seen in training.
I'm suggesting that something architecturally different is necessary to create a model which can make intuitive/comedic leaps.
Try to get a frontier model to write a clever, funny joke which hasn't been seen before. Or, try to get it to make an intuitive leap that leads to a novel discovery.
You can use them to guide your own efforts along these lines, as a sounding board. But with current architecture I just don't think either is possible for an LLM to do on its own.
ACCount37 1 days ago [-]
> Try to get a frontier model to write a clever, funny joke which hasn't been seen before.
This is what I refer to when it comes to larger models like Fable 5 being funnier. They are more capable of doing that. They can deliver that "sudden orthogonal leap from context" of yours more reliably.
It's not a "fundamental inability" and never was. If you crank the scale up and a capability appears, "current architecture" was never the problem.
nullbio 6 hours ago [-]
These capabilities that emerge are not "orthogonal leaps". They are very much just more of the same thing.
ACCount37 8 minutes ago [-]
Frankly, I doubt the existence of "orthogonal leaps" as a distinct category with distinct properties.
Just another emergent capability - one that follows exactly the same patterns any other emergent capability does.
RugnirViking 9 hours ago [-]
"there's currently no mechanism for producing an orthogonal directional change in the path"
idk the harnesses adding in "but wait -- " or some variation every few lines in the thinking trace seems to do great at this.
You posted this twice now. Why not just edit your original comment?
nativeit 20 hours ago [-]
I don’t need to edit my original comment, it doesn’t contain any mistakes. All six of the links across three comments are each distinct, separate posts going to either the same article, or others with identical headlines.
If you think LLMs can't "jump," watch Terence Tao engage with ChatGPT to try to understand the intellectual leaps achieved by Claude Fable 5 in finding a counterexample to the Jacobian conjecture https://chatgpt.com/share/6a5fdc7a-d6f8-83e8-bbea-8deb42cfed... .
After that you may revise your opinion.
GodelNumbering 1 days ago [-]
I have been writing a 'paper' [1] on an adjacent topic for months now. At some point, I decided to make it an empirical paper vs position paper. I am still chasing the experiments (when I get some free time waiting for agentic loops)
For this paper specifically, after reading the abstract [2], I felt almost certain that the author would have used Judea Pearl's ladder of causation (https://web.cs.ucla.edu/~kaoru/3-layer-causal-hierarchy.pdf) but they did not. Would have probably been a better argument to make.
[1] paper in quotes because it may never get published (it is over 20 pages atm). the core argument is that lack of native adjacency resolution makes problems harder and sample inefficient, not impossible
[2] "Using Einstein’s formulation of General Relativity as a case study, we demonstrate that LLMs are structurally incapable of creating new foundational axioms, particularly when observational data is scarce. "
Also, the claim that 'LLMs are structurally incapable of creating new foundational axioms' is provably false depending on where you place 'fundamental'.
JacobAsmuth 20 hours ago [-]
I've said this for a long time. This inability of AIs to generate novel explanatory hypotheses is a big blocker for their ability to automate jobs like accountant, middle manager, and even the lowly cashier.
ACCount37 1 days ago [-]
It's a very shaky position, and the empirical track record of "LLMs can't..." is in itself a reason to call it into doubt.
Every "can't" of this nature was followed by a discovery of "they can, just poorly", and then by that "poorly" improving steadily generation to generation.
The paper doesn't provide a way to measure or quantify this elusive "jumping" capability, not even as an approximation. It just throws "can't jump" out there, as if "abduction" is an established class of problem with known computational properties and requirements that the LLM architecture fails to satisfy. It's none of those things - and the paper makes the claim without backing it by anything but rhetoric attempts at persuasion.
The proposed solution is also dubious. The empirical track record of dedicated "world models" for reasoning and problem-solving is, frankly, downright abysmal. Even integrating multimodal data into LLMs has failed to yield general reasoning capability gains.
LeCun's misadventures in the field aside, the main frontier lab that pushes in favor of "improving reasoning via multimodal fusion" is GDM - and Gemini isn't exactly a paragon of frontier reasoning capabilities. It has strong multimodal capabilities, but lags behind both OpenAI and Anthropic in performance outside that - while Anthropic is the lab that always treated multimodal grounding as an afterthought, and still trades blows with OpenAI at the very edge of the performance frontier. Multimodal grounding seems to work great as a way to improve an AI's ability to deal with those specific modalities, but it falters outside that.
Now, it's not impossible that everyone who tried multimodal world models for reasoning is just doing it wrong, and there is an undiscovered recipe for multimodal grounding that results in a step change in AI capabilities. But the results we have so far suggest it to be unlikely.
vatsachak 1 days ago [-]
Okay here's something LLMs can't. They can't solve problems that are longer than ~10 pages of math. They also can't maintain codebases without supervision. It's because they are have no memory and use various tricks to supplant that fact.
ACCount37 1 days ago [-]
Now, how long before someone rolls out some sort of 10M context hybrid attention active context management monstrosity and ruins this guy's "can't"? Start the clock.
My opinion of claims like "LLMs need memory to manage codebases" has also hit the dumpster bin a while ago.
Why would knowing how to make a maintainable change to a codebase require any more "memory" than knowing how to play an optimal chess move? The codebase is the memory. A sufficiently capable LLM can ingest it, figure out what changes to make, and make them.
vatsachak 14 hours ago [-]
Okay then why do software engineers still exist. And if companies are making an error keeping people on instead of LLMs, why isn't there an all agent company making money off of it?
dandaka 10 hours ago [-]
1/ LLMs are doing increasingly bigger share of work of SWE with an increasing success rate
2/ SWE are doing way more than "writing code", and since those areas are less "computational verifiable" (what is a good architecture that will stand in 5y?) and more "connected to real world" (what are the requirements?), LLMs struggle with them
kskdkwkdjw 1 days ago [-]
> The codebase is the memory.
Were that the case LLMs would’ve been phenomenal code monkeys from the get go. They were not. They still are not.
petesergeant 1 days ago [-]
> the empirical track record of "LLMs can't..." is in itself a reason to call it into doubt ... Every "can't" of this nature was followed by a discovery of "they can, just poorly", and then etc
Maybe modulating temperature can help here: have the LLM come up with ideas at high temperature, and then critique them at low.
This is also tied to halucinations: it is something that humans do (for writing fiction, and for "jumps") - but what LLMs currently lack is intellectual honesty. Coming up with bullshit is fine (and in this context valuable) - the important bit is putting those ideas through some form of rigor, or just immediately turn around and admit to talking shit.
So I'd arge that hallucinations are what prevent LLMs from doing this in a useful way.
km3r 1 days ago [-]
I wonder if giving the models context of the temperature of its past generations would help here. Like a thinking mode that deliberately has a section that is high temperature, while the rest is lower.
zamalek 1 days ago [-]
I think you'd need to insert "critique these ideas." That's where the intellectual honesty comes in - I suspect that _all_ current LLMs would continue on with and truly stupid ideas as though they were gospel, their profound incompetence at backtracking on what they have said is something the industry hasn't figured out.
If we could make progress in that area, maybe CoT could gradually decrease as it approaches its limit, or maybe the LLM could control the temperature of the next token itself (how this would be trained, I have no idea).
1 days ago [-]
rsfern 1 days ago [-]
I found this paper really thought provoking, but I think the conclusion of “world models are the solution” leaves something to be desired. People are already equipping agentic systems with physical simulation tools and exploring action-conditioned world models. This is cool because you can change the rules of the simulation and observe what happens, but it doesn’t address the core question of what to change the rules to, or even what the goal should be in the first place.
1 days ago [-]
aantix 1 days ago [-]
Don't these datasets already exist in terms of the real-world samplings we have when training robots to do every day tasks?
Yopolo 1 days ago [-]
We just put different things together and then we evaluate it.
In math its simple: does the verification say its okay.
If its mechanical: is any property better than what we have already.
etc.
_superposition_ 1 days ago [-]
What I find absolutely fascinating about this paper is that recently I made the leap that physical representation was a necessary ingredient for invention based on my own experience (lack of abundance of evidence)
So I intuitively agree with the premise. It's kinda meta.
pama 1 days ago [-]
The early physics background is messy and incorrect. I didnt read the full position paper, but from its start: The Lorentz transformations were by Lorentz, well before Einstein’s paper on special relativity; the principle of relativity also existed before the Einstein paper. The math was all there, with steps taken by Maxwell, Voigt, Larmor, Lorentz, and Poincare. Einstein supplied a clean physical interpretation, making all inertial frames equivalent, making simultaneity frame dependent, and explaining length and time deformations without the need of the concept of ether. Skimming the end of the paper with the arguments about lack of abduction or inability to make the analogy without sensory experience, I see that this paper is unfounded speculation rather than solid/hard philosophical logic. As a position paper it is OK to appear, but i think it misses the point of how LLMs or other autoregressive learners of future states can build analogies and intuition that can help them formulate new theories of the world. Soon it will be more obvious to everyone, so I am not very worried about these writings.
skew-aberration 1 days ago [-]
Yes, I made a similar comment on the other mention today. The author seems to misunderstand what GR is / what it added to physics too (creating self-referential field equations to handle mass/energy equivalence - linear field equations without instantaneous 'action-at-a-distance' existed much earlier). The notion that it was a 'small signal' is totally false. Once you've hypothesized that the apparent mass of objects depends on the observer, you need to show that your theory gives consistent results for trajectories of objects in gravitational fields.
This tracks for me as someone making keeps of intuition in little-explored areas.
I just don't see any of the LLM users around at all. Clearly some force is guiding them all away from thinking any of the "leap of faith" thoughts that I am thinking.
dtj1123 1 days ago [-]
Neither can I, if I'm being honest.
brainless 1 days ago [-]
I have a weird thought experiment: If you give a GPT-2/3 level LLM tools to search the internet - any document, can it build bigger, better LLMs?
You may think this is not a good test because an older (or say a smaller) LLM can study from the knowledge on the Internet and build. But we are like that - we can access the Universe through our senses.
Can we ever produce anything that is beyond this Universe? I think an LLM that is lacking in knowledge can build more complex systems as long as it can access more data.
setnone 1 days ago [-]
LLMs don't have legs yet. If you're smart but can't touch things you only keep being smart
ACCount37 1 days ago [-]
[dead]
brcmthrowaway 18 hours ago [-]
Can talkie figure it out?
zie1ony 1 days ago [-]
Interestingly, halucinations might be the way to achieve that.
peter_d_sherman 1 days ago [-]
>"While Generative AI has mastered Induction (statistical pattern matching) and is rapidly conquering Deduction (formal proof), we argue it lacks the mechanism for Abduction—the generation of novel explanatory hypotheses."
Interesting!
Induction Vs. Deduction Vs. Abduction!
(You know, if you like Logic, Philosophy, Law, or... just plain different ways to think/reason about something! :-) )
yogthos 1 days ago [-]
There's no reason why LLMs can't be hooked up to a physical world feedback loop though, in fact that's already being done with chemistry research https://www.youtube.com/watch?v=AYSR02tcwes
m3kw9 1 days ago [-]
The proof is in the pudding, so far there isn't a proven (E=mc^2 type) breakthrough LLM's had made yet.
m3kw9 1 days ago [-]
Even myself, I really can't remember a time where I had this "jump". Is very subjective to feel this jump
I feel like you could just add some noise or randomness to the LLM and start approximating the leaps that the human mind uses to solve and understand unrelated things. Maybe that’s naive, it’s just coming from my organic computer in my skull.
dooglius 1 days ago [-]
Technically true, but the counter argument would be that the probability of this working would be ~ 2^(-(entropy_of_leap)) for an LLM (presumably intractable) and a human would succeed at a higher probability.
Mistletoe 1 days ago [-]
If we could just get the LLMs to take showers and dream, we would get some novel thoughts coming.
Der_Einzige 1 days ago [-]
Hahaha you just derived temperature from first principles.
Turns out temperature is pretty bad too, you can find ways to sample from deeper in the distribution without distorting it. Great example is XTC (exclude top choices), In a few weeks/months it'll also have a proper scholarly paper with peer review.
heaney-555 1 days ago [-]
So the goalposts have moved all the way to "LLMs can't do what Albert Einstein did"?
m3kw9 1 days ago [-]
Is this a reference to the movie "White man can't jump"? If so they are in a surprise because the movie says otherwise.
juleiie 1 days ago [-]
LLMs can’t but humans supplied with data and reasoning from an LLM can make novel jumps without absolute prior knowledge, or at least with reduced need for years of knowledge.
And then such jump can be verified by a machine so human kind of plugs the intelligence gap.
That’s pretty exciting.
I always liked to provocatively call LLM „the new calculator”. Calculator for language.
We are so focused on creating a standalone intelligence that we didn’t notice how we massively augmented our own. That could be considered transhumanism holy grail if only interface brain-LLM was faster.
People need to understand that these things are tools. And every tool needs an operator to function. Tool doesn’t have its own goals, needs, wants or motives. It won’t do anything out of its own, it always exists in context of someone telling it what to do.
In light of that most of the panic and fear mongering is rather ridiculous. Calculator won’t replace you. It wont take over the world. It is just a tool.
You write a book with book generator? Cool, it can be used for this. We will judge output, not the methods. Sometimes we will judge people who have no taste in literature.
prabhanjana_c 1 days ago [-]
Calculator for language, is a nice insight. I visualize generally as LLM's as big mathematical expression that produce next word.
But then , one difference I found on LLM's is on scaling laws, where at some point, it have interesting emerging properties, that nobody thought would be possible now being possible.
Tools:
Though It is a tool, but powerfull tool that is automating existing manually done jobs, large scale. Adapting to new roles where we definite goals, needs and motives judging of AI output, at large level is an issue. And humans are used to day to day repeating job, doing same thing repeatedly. Now the AI is taking over that. I see that is a challenge.
Also I see we are moving to creative world, where we will spend more time on creating something really new, leaving mechanized parts to AI.
1 days ago [-]
arklt 1 days ago [-]
[dead]
petsku 1 days ago [-]
[flagged]
shymaple 12 hours ago [-]
[flagged]
funflusion 1 days ago [-]
[dead]
luciana1u 1 days ago [-]
[dead]
slacker-gossip 1 days ago [-]
[dead]
reliablereason 1 days ago [-]
Clearly LLMs cant do leaps of intuition since their "intuition" is locked after training ends.
The only way a LLM can come up with new ideas if the "idea" appeared as a generalisation durring training or if it was achieved using reason in chain of thought.
Enginerrrd 7 hours ago [-]
I think this is a bad way to look at it. LLMs can probably conceive of most things that are representable within the embedding space.
Ordinarily in mathematics there’s a TON of papers to write just combining low level problems with different techniques. Better still, and often considered groundbreaking is borrowing techniques from other fields and adapting them or creating analogous methods to solve problems. A lot of landmark papers have been written this way. This is also what transformers are sort of good at within other contexts. They have super human breadth so I’m hopeful they’ll become real assets in math for a long time. Though the leaps necessary to adapt a technique in a nonobvious way might be too much for a while longer. We’ll see.
Truly novel techniques are quite rate indeed and I don’t know if LLMs can represent them faithfully in their embedding space or not. My inclination is that they probably can most of the time, but I don’t know. Mathematicians would describe such thins as “alien”.
buzzin__ 1 days ago [-]
Or some randomness is aomehowntroduced in the output, which happens after every word, unless you set the temperature to zero.
reliablereason 1 days ago [-]
That would not be intuition that is just randomness. Intuition is not randomness.
A jump in intuition comes from automatic processes reorganising the relational structure of conceptual models. There is no reorganisation of the model durring inference.
tpolm 1 days ago [-]
> There is no reorganisation of the model durring inference.
one could argue that model can reorganize / interact with prior knowledge captured in text form (edit files) hence there can be reorganisation
So, I do sort of buy into this idea that Einstein was simulating the world and running experiments on those simulations in ways that were beyond what you could encode in natural language. Will AI be capable of doing this, if it is bounded by training data that is composed almost entirely on language? One might argue that if AI is training on a lossy encoding/representation of the human experience, how will it be able to simulate anything beyond that experience? Unless it does so in a way that we manage to do when we image objects beyond 3D. But now I'm just rambling.
With apologies if this is common knowledge at this point, 3Blue1Brown has been doing an excellent series on compression, and its relationship to intelligence (or more controversially, that they are one and the same): https://www.youtube.com/watch?v=l6DKRf-fAAM
But that also throws in sharp relief, that there is vastly more to the human experience than intelligence alone: qualia, desire, gut instinct, intuition, emergent creativity. (Whether the "God of the Gaps" for the delta between capabilities of human vs AIs is fixed, or diminishing, or even shrinking to zero, remains an open experiment we're all living through.)
The compression series sounds interesting!
I think the deltas are growing at different rates.
The 3b1b videos mention how this concept may be used to find similarities between different languages. Researchers have been able to get results that closely resemble how languages are usually grouped into families.
Math is pure reasoning and stock LLM is already better at it than 99.9% of humans.
And as for human language, agentic LLM can produce a perfectly human text. The fact that AI texts have tells is just because default settings are the same for millions of people using given LLM and almost everybody just pretty much on-shots the text instead of doing it agentically with anti-slop detection and rephrasing in the loop.
Yes, and sometimes this is very intentional. Take for example a short poem which if you sit and really think about it for a long time, you could go off on a mental tangent of imagining what sort of kingdom or empire created a statue that is now "two vast and trunkless legs of stone", for instance. Being terse and allowing for human interpretation is kind of the entire point of something being written like this.
I met a traveller from an antique land
Who said: Two vast and trunkless legs of stone
Stand in the desert. Near them, on the sand,
Half sunk, a shattered visage lies, whose frown,
And wrinkled lip, and sneer of cold command,
Tell that its sculptor well those passions read
Which yet survive, stamped on these lifeless things,
The hand that mocked them and the heart that fed:
And on the pedestal these words appear:
"My name is Ozymandias, king of kings:
Look on my works, ye Mighty, and despair!"
Nothing beside remains. Round the decay
Of that colossal wreck, boundless and bare
The lone and level sands stretch far away.
It's also to a certain extent why LLMs work.
When IBM Watson was playing Jeopardy, one of the game prompts was:
> It was the anatomical oddity of U.S. gymnast George Eyser, who won a gold medal on the parallel bars in 1904
The man was missing a leg and used a prosthetic. Watson's output was, "What is leg?"
At first it was regarded as correct. If a human said that you could conclude that they knew the answer. But then the judges decided not to give Watson the point because its output didn't provide enough specificity to prove that it understood the context.
If you ask an LLM what kinds of things taste sweet it can give you examples like cotton candy or strawberries, but it has never actually tasted anything. All it knows is that the training data contains the association between those tokens. But the human reading the output knows what strawberries are, which is what allows the output to be meaningful.
> The fundamental problem of communication is that of reproducing at one point either exactly or approximately a message selected at another point. Frequently the messages have meaning; that is they refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem.
"A post-graduate student equipped with honours and diplomas went to Agassiz to receive the final and finishing touches. The great man offered him a small fish and told him to describe it. Post-Graduate Student: “That’s only a sun-fish” Agassiz: “I know that. Write a description of it.” After a few minutes the student returned with the description of the Ichthus Heliodiplodokus, or whatever term is used to conceal the common sunfish from vulgar knowledge, family of Heliichterinkus, etc., as found in textbooks of the subject. Agassiz again told the student to describe the fish. The student produced a four-page essay. Agassiz then told him to look at the fish. At the end of the three weeks the fish was in an advanced state of decomposition, but the student knew something about it."
sourced from: https://nabeelqu.co/understanding
—Albert Einstein
Quoted in Using Spaced Repetition Systems to See Through a Piece of Mathematics,
https://news.ycombinator.com/item?id=18895613
which describes the author's experience that if you approach a field obsessively enough, eventually you begin to understand it at a level deeper than language.
If an LLM is big enough, I imagine something similar is happening.
Then again maybe the philosophy department has a different opinion :)
If you read every single book there is about The Grand Canyon, and watched every single video and/or documentary about The Grand Canyon, do you believe that you have fully experienced The Grand Canyon? Or do you just have to be there to fully experience it.
I dunno. Substitute in whatever you want for "The Grand Canyon". Maybe climbing Mount Everest or walking on the Moon. The point is that maybe the human experience is more vast than what is written about it.
But in terms of objective knowledge about the Grand Canyon, reading every article and scholarly work would give you a much better grounding. Unless you're a geologist/ecologist doing actual field research there, actually seeing the Grand Canyon would add little, if anything, to your factual knowledge and understanding of the site.
If I'm going to hire a guide, I'm going to hire the guy who spent a week hiking it instead of the guy who spent a week reading about it or watching videos about it.
Rather one should ask "Would you be able to answer any question and fulfill any request about the Grand Canyon given to you by others in a way that will be indistinguishable from someone who has been to the Grand Canyon"
You might equally contrast someone who is physically in the Grand Canyon with an LLM given access to a drone with sensory attachments. I'd bet that if you told groups of humans and drone-controlling AIs to find something interesting, the humans would be more variable and cover a wider range of interesting discoveries but often just give up on the task, while the AI would be more likely to find something but would cluster around certain discoveries and make fewer overall.
I'd be fantastically surprised if you could provide a precise answer to any such question without extensive security footage that doesn't exist.
An estimate, perhaps, but you probably couldn't even answer a simpler "how many people are currently in the canyon" without resorting to estimation given the lack of information currently available via the mentioned sources.
Also, you touched on how even modern tech still falls short of the true experience. Just look at the history of motion picture since the early 1900s. We have added sound, color, bigger screens, faster refresh rates, more pixels, 3D (sort of), surround sound, spatial audio, IMAX, etc. Almost seems like video leaves a lot to be desired.
Worth... quite a lot imo.
First, in the slightly objective sense of "What is this place like in real life, under X conditions." But, more subjectively, watching a video of a glacier and imagining the wind/rain/cold doesn't even approach 5% of the intensity and awe of climbing up a mountain yourself to see a seemingly endless expanse of ice, struggling to stand steady because of the wind, shivering because of the cold and rain. And finally, in the financial sense, a lot of people routinely spend many thousands of dollars and days-weeks of their time to experience natural wonders in real life.
It's not that the value of the grand canyon is low in an absolute sense, it's that in a relative sense the gap between "empty numb void" and "hiking at home (plus grand canyon videos)" is far far greater than the gap between "hiking at home (plus grand canyon videos)" and "actual grand canyon".
Sure, no doubt. But I don't really think "video of my local state park + imagining the parts that aren't conveyed over video" is particularly close to the real thing either.
Have you ever been to a place which is really high up? You're above the trees and can not only see as far as the curvature of the earth allows you to, but are high enough that the curvature of the earth allows you to see for miles. The air is thinner so it's easier to take a breath yet harder to catch it. Your sweat evaporates faster but a breeze doesn't cool you as much. The sun is brighter, rises earlier and sets later. The wildlife makes unfamiliar sounds and sound itself is changed. Even the same food tastes different.
It's not something you can get by sitting at home and changing the background on your computer.
Back to the topic, I'm sure it's very impressive, but I think you're underestimating just how impressive our baseline is.
It's similar in dismissing artificial audio based on listening to some pop song on a basic mobile speaker versus listening to well-designed binaural audio on good headphones. The latter can sound 'freakily' 'real'.
But ultimately this is just a matter of training data. I do not say that it is easy to obtain the required data, but it is not a fundamental problem LLMs can't overcome.
https://www.noahpinion.blog/p/what-will-more-intelligence-ac...
> Another way of saying this is that there may be laws of the universe that humans can’t understand but AI can. I call these “cloud laws” — causal regularities that can be exploited by technology, but which are too diffuse and complex for an individual human being to either intuit or communicate. Human language seems to obey cloud laws, so why not other phenomena too? Perhaps social sciences like economics, sociology, and political science obey similarly complex regularities, and AI can help us find them. Perhaps there are physical processes — plasma, or topological materials, or aerial turbulence, etc. — that obey cloud laws instead of chaos?
Building AGI Using Language Models – https://bmk.sh/2020/08/17/Building-AGI-Using-Language-Models...
> From the two postulates, Einstein derived the Lorentz trans- formation ...
If Einstein derived them, who is "Lorentz"?
The groundwork for Special Relativity was the study of electrodynamics and symmetries of Maxwell equations. The Einsteins paper was literally called "On the Electrodynamics of Moving Bodies" and never cites Michelson and Morley.
Einstein is not the first great scientist who are in denial of other important prior contributions, and he also not the last one. Newton also probably knew too well about Al-Haytham (Alhazen), arguably the father of modern science, and his breakthrough experiments but never directly cited Alhazen's works in his seminal books on Optics.
[1] Millikan, Einstein, and the Birth of Relativity (4 letters):
https://www.aps.org/archives/publications/apsnews/200403/let...
"Abraham Pais, who knew Einstein well and wrote his scientific biography, was certain that Einstein did know about Michelson's experiment before 1905. He points out that Einstein was over seventy and in poor health when he spoke to Shankland; at the first interview he probably did not remember that Michelson's experiment is discussed in Lorentz's 1895 monograph, the famous "Versuch", which Einstein had definitely read before 1905."
Furthermore I can't help but feel some logic would lead us to conclude it is not natural to cite, why? Because in our time if you do not cite your sources in a lot of occupations dealing with knowledge you will be punished for it. It would be absurd to construct punishments for failing to do what people will naturally do, and to have those punishments so often called on.
Even if we give a very lenient interpretation of not citing sources we would probably have around 10% of papers with errors, and that is with the fear of punishment.
Citing when one should and accurately is not normal, it requires discipline to do it, or there would not be such a high failure rate at doing it.
So it seems you have established that no, Newton was not expected to cite everyone in his time, very well, this now puts the question further up the chain over my entry into the conversation - "Were the people Newton neglected to cite actually integral to his work?"
My definition of integral would be, two possible definitions:
1. do I write something here that someone reading it would say, "What?! What do you mean, explain yourself, how did you arrive at this surprising bit of information?!!" then it is integral.
2. Am I doing this based on some work that someone did without which I would definitely not be doing this? If I am doing it with knowledge of the prior work but whether or not the prior work existed I would still be doing it, the prior work is not actually integral.
[0] https://xkcd.com/1053/
https://arxiv.org/abs/0908.1545
Science rightfully recognizes the mind that doesn't just first describe the idea but provides a robust framework to test it and communicate it.
> If Einstein derived them, who is "Lorentz"?
You can (re-) derive a lot of existing stuff.
Einstein was aware of Lorentz and the transform. He was aware of Poincaré as well. He knew the state of the art for his time.
Imho there's also a clean argument against the existence of LLM-understanding: chatbots have been unable to summarize to experts (see my reply to you in the other thread) their own findings.
Even after prolonged interrogation. They were unable to _compress_ their own findings. Thus they might not actually understand what they have actually done. (They might barely pass an oral thesis defense)
(One may object-- that proofs aren't data that can be "compressed". But then doesn't the process of abduction generalise the very idea of data? to.. ?)
But I'd like more explanation as to why Maxwell or Newton are any less replaceable: all 4 of "Maxwell's equations" have other names attached to them (the exception is Ampere-Maxwell law [1] which Maxwell contributed an important term to), and Newton had a number of contemporaries who were making similar discoveries but are often forgotten.
[1]: https://en.wikipedia.org/wiki/Amp%C3%A8re's_circuital_law
The size of an atom was computed for the first time 40 years earlier, in 1865, by Johann Josef Loschmidt, who was an Austrian scientist.
Loschmidt had determined the value of what today is called "Avogadro's number", despite the fact that Avogadro had no idea about how one could find the value at that number. The contribution of Avogadro had been the law that ideal gases have the same number of molecules per volume when the pressure and the temperature are the same, but he did not know how that number might be determined.
Based on the results of Loschmidt, which were also used by Maxwell in his works on statistical physics, a few years later, in 1874, George Johnstone Stoney computed the value of the elementary electric charge, which he later named "electron".
Loschmidt was the one who first measured and weighed the atoms, establishing their existence beyond any reasonable doubt. Before him, many believed that the hypothesis of Dalton about the existence of atoms is just a convenient theory for explaining the laws of chemistry, but some other better explanation might be discovered in the future.
Einstein published an explanation of the Brownian motion, which verified that everything fits in the known theories and it matches the expected values. It did not contain any novel number or law, but it was a useful confirmation of the current theories.
In my opinion, the greatest achievement of Einstein has been the paper published by him in which he revealed the existence of the stimulated emission of electromagnetic waves by hot bodies, besides their well known absorption and spontaneous emission.
This paper had tremendous practical applications, leading to the discovery of masers and lasers, without which the modern technology would have never existed.
While Einstein's theory of gravitation was more intellectually challenging, its practical applications are minor, almost negligible. Moreover, even after more than a century we do not know exactly how adequate it is for describing gravitation.
What is certain is that Einstein's theory of gravitation cannot be anything else but an approximation, as it is based on averaging the distribution of matter in space, without taking into account the concentration of matter into elementary particles.
Another contribution of Einstein that is more important than his most frequently cited works is his support for Satyendra Nath Bose, which lead to the recognition of the Bose-Einstein statistics, which together with the Fermi-Dirac statistics plays an essential role in quantum physics.
What I have written is not a dismissal of Einstein. On the contrary, I have said that Einstein is the source of a few extremely important discoveries, but those are not the theories cited by most people, who have never read the original work of Einstein but speak about him from hearsay.
As I have written above, when speaking about Einstein, almost nobody mentions the discovery of stimulated emission, which has been far more important than anything else done by Einstein. Einstein did not discover special relativity, the photoelectric effect or Brownian motion. He just provided alternative explanations for them, which did not include any new formula or quantity.
On the other hand, stimulated emission was a surprise, even if it is one of those things that appears to be obvious in retrospect, but for decades nobody who studied the blackbody radiation had noticed that something is missing in the equations.
This discovery eventually lead to the invention of masers and lasers and today there exists almost nothing that incorporates modern technology and which could exist without lasers.
Already for decades, no modern integrated circuits, no CPU and no memory can be fabricated without using lasers. Even many purely mechanical things, but which require great precision, cannot be fabricated without using measurement instruments that depend on lasers.
So the contributions of Einstein are very important, but they are not those that most people associate with him. Besides stimulated emission his most original work was his theory of gravitation, a.k.a. general relativity, but if he had not developed it that might have changed nothing, because Hilbert had developed an equivalent theory almost simultaneously, so had Einstein not published first, Hilbert would have become its discoverer.
After stimulated emission, the most important in practice research of Einstein is the Bose-Einstein statistics. But again, here like for special relativity, the statistics had been discovered by someone else, i.e. by Bose. Einstein has nonetheless the great merit of understanding its value. He demonstrated that the statistics was much more generally applicable than initially believed by Bose and his authority convinced everybody about its importance.
Thus there is no doubt that Einstein is one of the most important physicists of the 20th century. Nevertheless, there are dozens who made similarly important discoveries, e.g. Sommerfeld, Dirac, de Broglie, Schroedinger, Fermi, Rutherford and many others.
http://news.bbc.co.uk/2/hi/science/nature/541840.stm
Einstein is indeed one of the most important physicists of the 20th century, but it is really meaningless to try to establish a hierarchy between physicists, because their works are like open-source programs, each work is based on the works of the predecessors and it would have been impossible to do without those previous works.
In any case, Einstein was good, but Maxwell seems really superior in the sense that he could extract the essential information from the much more confuse sources that were available at his time.
The Treatise of Electricity of Magnetism of Maxwell and his collected papers are very instructive lectures even now and there are fragments of them that remain better than the equivalent fragments in many modern textbooks, despite the huge additional information that we have today.
Most of the original texts written by Einstein are now of interest only for understanding the details of how his thinking and that of the contemporaneous physicists have evolved during those important years, but those texts have little other importance except for history, because later better expositions of the same material have appeared. Nonetheless, it is important to also read Einstein's originals, because in many later books there are some inappropriate interpretations attributed to Einstein, but in reality Einstein had said something else and which was more correct than what some later author misunderstood.
fair, there's lots of really awesome scientists, many of which not talked about even 1/100th as often as einstein
>when he doesn’t even belong in that conversation.
suddenly, the pendulum has swung way too far in the other direction.
What a depressingly peculiar time to be alive.
Reddit is over that way, my friend. You might find the crowd there more amenable to this nonsense.
Einstein’s fame isn’t the problem. It’s the narrative that he is somehow the greatest scientist ever, when it’s just transparently false when you look at his actual contributions and the historical momentum of the fields he contributed to. There’s nothing wrong with his fame, the problem is that it overshadows actual juggernauts, like Maxwell and Newton. Maxwell not being famous at all among the ordinary public is a great tragedy, when he genuinely is in the conversation as having have been the most important physicist to have ever lived.
> A few reflections on my "LLMs Can’t Jump" paper:
> My position paper recently got some traction here, so I wanted to share a few thoughts and clarify a few things.
> First things first: some people are framing this as "DeepMind is throwing cold water on AI for science" or claiming the paper argues LLMs can never make real scientific discoveries. This is NOT the case.
> This is a personal position paper, not the company's view on AI for science. This is also not my position. As a core contributor to AlphaProof (the first AI system to win an IMO medal), I know firsthand that my colleagues at DeepMind, other frontier labs, and academia have made amazing discoveries with LLMs and will continue to do so. This paper is NOT an "LLMs are a dead end" kind of thing.
> Rather, the paper is the result of a deep dive I took to study the invention of General Relativity. I wanted to explore what it would take for a modern AI system to make that exact kind of jump. Specifically, I focused on the equivalence principle—a key axiom that Einstein formulated through thought experiments grounded in his physical intuition. I was trying to figure out what it would take to give modern AI systems that sort of thinking.
> Giving AI this specific capability isn't necessarily the most urgent thing to do next. It is very likely that improving our current recipes will lead to many exciting discoveries in the near future. In fact, that is what I am personally working on these days (sorry to disappoint you!). It is also quite possible that I am wrong, and that simply scaling our current systems will lead to new inventions in physics and elsewhere.
> Nevertheless, this was my position last winter when I wrote the paper, and I'm sticking to it. I think that there are a few interesting ideas to explore in this space which could influence the next generation of AI systems. I was very lucky to receive a lot of interesting feedback about this position—thank you for all the messages!
[1] https://x.com/TZahavy/status/2082401499628376180
It's weird because the equivalence principle is very unintuitive. Aristotle's Mechanics does not have it. It took almost two thousand years to discover inertia that is the most simple version of the equivalence principle. Einstein understood the idea of the the equivalence principle because he had a physics degree, not because he feel that in real life.
Moreover, if you ever have to study or teach Quantum Mechanics, physical intuition gets in the way. A lot of properties contradict the physical intuition but after a while you get use to them. If we continue with Einstein, the photoelectric effect does not aperar in real life.
Of course it does. How do you think your phone camera works?
Another bad idea for a movie is a watermillpunk universe, where during a practice on a hot day the best ever curling player discover inercia.
https://en.wikipedia.org/wiki/Oil_drop_experiment
(Note that the HP Chipmunk 9836 also had a scroll selector wheel in 84)
Most people are UI bigots. Once they get used to a first something, they expect everything to work that way and hate learning anew. They get stuck on keyboards, mice, trackpoint nubs, trackpads, trackballs, scroll wheels, or touchscreens and refuse to move on. Of course there are 'objective' performance tests for each including Fitt's test of accuracy and latency, as well as, cognitive load. I guess once you have a hammer, every screw looks like a nail.
So I'd be a bit careful with the "of course!". It may well be obvious only because that is the first mental model you latch onto.
1. Read the last sentence of the abstract, and
2. Reflect that frontier reasoning agents already increasingly integrate multimodal models.
I simply point out here, that fully accepting the paper’s premise, the paper’s conclusion isn’t limiting on frontier AI reasoning agents. The paper posits the necessity of multimodal world models and the limitations of LLMs. Frontier agents aren’t simply LLMs and do increasingly integrate increasingly capable multimodal models.
> Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
> Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.
> When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
TFA was actually about leaps of intuition, sadly.
One of the experiments I've heard proposed around here is to somehow create an LLM from all text up to 1980 or 1990 and see if it can get back to making itself.
Could be, but preventing leakage from more modern stuff can be challenging.
This was attempted with Victorian public domain content: https://www.estragon.news/mr-chatterbox-or-the-modern-promet...
I can't find the citation right now, but I think people found it was leaking anachronisms? So this probably wasn't as well filtered as the creator had hoped?
At a minimum, yes. IIRC, the sum total of all compute manufactured over history only reached the minimum needed to train an OK LMM in the mid 00s.
> How much could it infer from it?
Only way to find out is to try.
https://didof.dev/blog/talkie-1930-llm-reasoning/ seems like a decent overview
This is an interesting experiment but I wonder if it would be possible to prevent some sort of retrospective bias. For example, I’d expect the experiments that lead to relativity to be over-represented in our catalogue of scientific literature prior to 1900, just because in retrospect they were important, so the records about them were preserved.It would have to be a very intentionally constructed corpus, I think.
IMO, the mechanism isn't the important thing, the behaviour is. If you look at the step-by-step, we are also looking for the next word or motor action (and for whoever is about to suggest that we humans plan ahead, Transformer-based LLMs have been shown to also do this); as this is not a useful description of what it means to be a living brain, I'd say it's also not a useful description of what makes everything post-InstructGPT different from what came before.
I know the answer: because it leads to model collapse. But why is that? Wouldn't a smart model not collapse? It's seeming like they keep getting smarter because we keep pouring more of our own knowledge into them, not because they are actually getting smarter. And yes, sometimes a dumb but persistent bruteforcer can make new discoveries.
and i think this is exactly the crux;
the really big models need really big datasets
and current gen LLMs get a lot of training data beyond "all books + all of the internet"
the objection is then that producing this additional data would already confound it with pre "virtual cutoff date" knowledge (since the training data probably implies mathematical and SWE concepts that were developed post "virtual cutoff date")
But to prevent model collapse you need a way to pump down the entropy. Much like in thermo, it's an expensive and slow process.
My feeling is that a prompt would have to provide a vague description of a program that meaningfully passes something like a Turing test, an API to conform to, an expectation of novel construction (no 'ifs all the way down'), and then a requirement to search broadly and pursue promising ideas and not get hung up on the philosophy. Anything more precise feels like it would corrupt the test, but as it is that description feels doomed to loop before even trying the interesting parts.
Chessboxing was the invention of comics book artist Enki Bilal (and he's credited with this in Wikipedia). I first saw it in his Nikopol trilogy. Because life is weird, it then became a real thing.
It's unrelated to computers playing chess. It predates Kasparov's first defeat by Deep Blue. I don't remember any mention of computers being good at chess in the trilogy, either. Or any computers, for that matter.
It's actually possible to answer this question rigorously:
1. Define a scientific result which qualifies as a "jump". They should be frequent enough that they happen every year - otherwise one might say humans can't jump either.
2. Identify all such "jumps" in articles published in 2026, and use LLM with 2025 knowledge cut-off to re-derive these results with minimal amount of information.
It really irks me that people boost these low-effort articles just because they confirm pre-conceived notion that LLMs are limited
To Yann LeCun it was "obvious" that you can't build a world model from a text, so he predicted that even "GPT 5000" can't predict that an object sitting on the table would move together with the table. But GPT-4 could solve this task and provided correct explanation. GPT-5 can solve much more complex tasks, code physics simulation, etc.
Now people talk about invention of General Relativity as something LLMs can't do... something almost no human was able to do either.
"In abstract domains such as Mathematics or Computer Science, the Sense Experience (E) may be grounded in high-dimensional topology or have other goals such as generality or minimality."
But if such sense experience is possible in abstract domains via some high-dimensional topology, why could a sufficiently advanced LLM not develop an equivalent high-dimensional topology for domains like physics and use it to make creative leaps?
But we have no idea at how good humans are at that. Given the appalling failures of humans to handle even basic statistical situations like identifying that the same thing happens over and over, it might be that they are hilariously bad at creative leaps in abstract fields, it is just we have had nothing better available to measure against. We've spent about as long as decision theory existed trying to convince people to use it instead of flailing. Limited success, usually in exceptional cases.
And the paper seems a bit dodgy, we have models created with sensory data available. No reason a LLM can't be trained on more sensory data than a human can accumulate in one lifetime. There is a lot of visual data on YouTube.
The case study they chose is literally Albert Einstein coming up with General Relativity, something most scientists of his time were not able to do
One interesting (albeit sad) area which might be related are humans who are never raised with a first language. They seem to never developer abstract reasoning and even seem to lose the ability to develop it later in life. This might indicate there is some 'real world senses' -> 'direct language' -> 'indirect language' -> 'abstract abduction' hierarchy that develops, perhaps related to more real world abductions as a necessary side chain to developing abstract ones.
One of the obvious problems with this is just how difficult we find it to study intelligence purely in humans. We are measure a LLMs by a yardstick that is already known broken, but maybe this is still the right path.
Until LLMs have some 0% error humans will have to be in the loop (even if they only serve to take responsibility of the process).
Also possible if you make god
Saying you want to make workers more productive and provide better tools for people just isn't that sexy.
That said, here's an experiment conducted by some Soviet psychologist, I forgot the name. The man wanted to study intuition. So he invented an experiment that was supposed to trigger it in laboratory conditions. (Take a moment to marvel at that; how would you approach such a task?) He gave people a few puzzles. One was to place some sticks according to some rules. Yet another was to find a path in a maze. The secret was that the path in the maze was the same figure as the solution to the stick puzzle.
And he observed interesting results. People who solved the maze after the sticks found the path much faster than the control group. If a subject was asked to comment how he was solving the maze, at the start or halfway through, the speed dropped to typical. Subjects normally didn't notice the similarities.
So there is something to study here, although it is obviously a case of pattern matching, only subconscious. This is a jump of sorts, but not the one I mean. What I mean is a Zen jump.
An LLM in isolation from its environment might as well be a brain in a vat in some dark cave. You need an external environment to sample from and act upon to make forward progress.
I see similar thinking in stories of how humanity got here. Religion has thousands of years adapting to this problem, every time we explain something, the goal post moves. Catholics today accept evolution (or least the church does), but it is the "jump" from monkeys to humans where God is the only explanation.
Just 5 years ago we didn't have a technology that knows more about everything than even most experts. We keep coming up with benchmark after benchmark and LLM/AI keeps destroying them. Now we've moved the benchmark to "the jump". Again, maybe it's LLMs or the way we currently do them that can't do this, but eventually something will.
We still don't have that. LLMs have shown time and time again that they don't know a single thing and are incapable of reasoning.
And we still don’t. What we have are simply very advanced search results aggregators with delusions of personality. Just because your fridge says “I” doesn’t mean it is a person.
For example "..ARC captures the logical leap, it misses the manipulative component—the physical sensation and embodied simulation..." makes lots of assumptions on how such a discovery must occur, e.g. through "physical sensation and embodied simulation". Results matter, not the path there.
For example, quantization of energy, at the core of QM, wasn't discovered through "physical sensation and embodied simulation" at all. Planck simply found that if energy is quantized, then one obtained the observed black-body radiation spectrum. There was no "physical sensation and embodied simulation".
LLMs are inherently probabilistic, and there's currently no mechanism for producing an orthogonal directional change in the path traced through a latent space which is also contextually relevant (landing on a punch line).
In other words, LLMs are fundamentally incapable of making intuitive/orthogonal leaps in context.
It might be possible to add this capability with a new architectural component like transformers, but specifically for making "left turns"/intuitive leaps.
[1] https://wnmurphy.com/llms-cant-do-humor/
"Probabilistic next word prediction" and "humor" sit about as far apart as "modulating airflow with meat flaps" and "humor" do. One is an interface through which an action is performed and the other is a highly abstract capability.
Would you claim that a podcast comedian is fundamentally incapable of being funny because all he ever does is wiggle the air with his throat meat flaps? Probably not.
Absolutely nothing about "probabilistic next word prediction" forbids "making intuitive/orthogonal leaps in context". The interface is expressive enough.
And empirically? The "sense of humor" in LLMs is yet another "a function of model scale" capability. GPT-4.5 was reportedly funnier than both GPT-4o and o1. Fable 5 is reportedly funnier than Opus 4.x. It's one of those ever-elusive "big model smell" signs that are hard to measure with anything other than vibes.
Under the "humor as an opposed social intelligence test" family of hypothesis, what "being funny" reflects is the funny guy's ability to model and predict you and your reactions. For the comedian to be able to make the audience laugh, he must know his audience well, model it accurately enough to be able to spot the "breaking points" of humor, things they'd find unexpected and clever and thus "funny", and then weave those things into the jokes.
Then, a bigger LLM gets better at humor because it has a more accurate model of how humans think of things - including the "ha-ha" gaps. It's a "theory of mind" capability. It's not "special", it's just hard.
I think you're saying that you can eventually train models to arrive at that destination by training on existing jokes, effectively encoding these leaps as probabilities.
In that case, the model isn't actually making an intuitive/comedic leap; they're just following new probability chains in attempting to approximate examples they've seen in training.
I'm suggesting that something architecturally different is necessary to create a model which can make intuitive/comedic leaps.
Try to get a frontier model to write a clever, funny joke which hasn't been seen before. Or, try to get it to make an intuitive leap that leads to a novel discovery.
You can use them to guide your own efforts along these lines, as a sounding board. But with current architecture I just don't think either is possible for an LLM to do on its own.
This is what I refer to when it comes to larger models like Fable 5 being funnier. They are more capable of doing that. They can deliver that "sudden orthogonal leap from context" of yours more reliably.
It's not a "fundamental inability" and never was. If you crank the scale up and a capability appears, "current architecture" was never the problem.
Just another emergent capability - one that follows exactly the same patterns any other emergent capability does.
idk the harnesses adding in "but wait -- " or some variation every few lines in the thinking trace seems to do great at this.
https://news.ycombinator.com/item?id=49136070 https://news.ycombinator.com/item?id=49096837 https://news.ycombinator.com/item?id=46890333 https://news.ycombinator.com/item?id=46870562
All of these are titled “LLMs Can’t Jump”
For this paper specifically, after reading the abstract [2], I felt almost certain that the author would have used Judea Pearl's ladder of causation (https://web.cs.ucla.edu/~kaoru/3-layer-causal-hierarchy.pdf) but they did not. Would have probably been a better argument to make.
[1] paper in quotes because it may never get published (it is over 20 pages atm). the core argument is that lack of native adjacency resolution makes problems harder and sample inefficient, not impossible
[2] "Using Einstein’s formulation of General Relativity as a case study, we demonstrate that LLMs are structurally incapable of creating new foundational axioms, particularly when observational data is scarce. "
Also, the claim that 'LLMs are structurally incapable of creating new foundational axioms' is provably false depending on where you place 'fundamental'.
Every "can't" of this nature was followed by a discovery of "they can, just poorly", and then by that "poorly" improving steadily generation to generation.
The paper doesn't provide a way to measure or quantify this elusive "jumping" capability, not even as an approximation. It just throws "can't jump" out there, as if "abduction" is an established class of problem with known computational properties and requirements that the LLM architecture fails to satisfy. It's none of those things - and the paper makes the claim without backing it by anything but rhetoric attempts at persuasion.
The proposed solution is also dubious. The empirical track record of dedicated "world models" for reasoning and problem-solving is, frankly, downright abysmal. Even integrating multimodal data into LLMs has failed to yield general reasoning capability gains.
LeCun's misadventures in the field aside, the main frontier lab that pushes in favor of "improving reasoning via multimodal fusion" is GDM - and Gemini isn't exactly a paragon of frontier reasoning capabilities. It has strong multimodal capabilities, but lags behind both OpenAI and Anthropic in performance outside that - while Anthropic is the lab that always treated multimodal grounding as an afterthought, and still trades blows with OpenAI at the very edge of the performance frontier. Multimodal grounding seems to work great as a way to improve an AI's ability to deal with those specific modalities, but it falters outside that.
Now, it's not impossible that everyone who tried multimodal world models for reasoning is just doing it wrong, and there is an undiscovered recipe for multimodal grounding that results in a step change in AI capabilities. But the results we have so far suggest it to be unlikely.
My opinion of claims like "LLMs need memory to manage codebases" has also hit the dumpster bin a while ago.
Why would knowing how to make a maintainable change to a codebase require any more "memory" than knowing how to play an optimal chess move? The codebase is the memory. A sufficiently capable LLM can ingest it, figure out what changes to make, and make them.
2/ SWE are doing way more than "writing code", and since those areas are less "computational verifiable" (what is a good architecture that will stand in 5y?) and more "connected to real world" (what are the requirements?), LLMs struggle with them
Were that the case LLMs would’ve been phenomenal code monkeys from the get go. They were not. They still are not.
What LLMs are "fundamentally incapable" of doing has striking parallels to https://en.wikipedia.org/wiki/God_of_the_gaps
This is also tied to halucinations: it is something that humans do (for writing fiction, and for "jumps") - but what LLMs currently lack is intellectual honesty. Coming up with bullshit is fine (and in this context valuable) - the important bit is putting those ideas through some form of rigor, or just immediately turn around and admit to talking shit.
So I'd arge that hallucinations are what prevent LLMs from doing this in a useful way.
If we could make progress in that area, maybe CoT could gradually decrease as it approaches its limit, or maybe the LLM could control the temperature of the next token itself (how this would be trained, I have no idea).
In math its simple: does the verification say its okay.
If its mechanical: is any property better than what we have already.
etc.
https://news.ycombinator.com/item?id=49177965
I just don't see any of the LLM users around at all. Clearly some force is guiding them all away from thinking any of the "leap of faith" thoughts that I am thinking.
You may think this is not a good test because an older (or say a smaller) LLM can study from the knowledge on the Internet and build. But we are like that - we can access the Universe through our senses.
Can we ever produce anything that is beyond this Universe? I think an LLM that is lacking in knowledge can build more complex systems as long as it can access more data.
Interesting!
Induction Vs. Deduction Vs. Abduction!
(You know, if you like Logic, Philosophy, Law, or... just plain different ways to think/reason about something! :-) )
Turns out temperature is pretty bad too, you can find ways to sample from deeper in the distribution without distorting it. Great example is XTC (exclude top choices), In a few weeks/months it'll also have a proper scholarly paper with peer review.
And then such jump can be verified by a machine so human kind of plugs the intelligence gap.
That’s pretty exciting.
I always liked to provocatively call LLM „the new calculator”. Calculator for language.
We are so focused on creating a standalone intelligence that we didn’t notice how we massively augmented our own. That could be considered transhumanism holy grail if only interface brain-LLM was faster.
People need to understand that these things are tools. And every tool needs an operator to function. Tool doesn’t have its own goals, needs, wants or motives. It won’t do anything out of its own, it always exists in context of someone telling it what to do.
In light of that most of the panic and fear mongering is rather ridiculous. Calculator won’t replace you. It wont take over the world. It is just a tool.
You write a book with book generator? Cool, it can be used for this. We will judge output, not the methods. Sometimes we will judge people who have no taste in literature.
But then , one difference I found on LLM's is on scaling laws, where at some point, it have interesting emerging properties, that nobody thought would be possible now being possible.
Tools: Though It is a tool, but powerfull tool that is automating existing manually done jobs, large scale. Adapting to new roles where we definite goals, needs and motives judging of AI output, at large level is an issue. And humans are used to day to day repeating job, doing same thing repeatedly. Now the AI is taking over that. I see that is a challenge.
Also I see we are moving to creative world, where we will spend more time on creating something really new, leaving mechanized parts to AI.
The only way a LLM can come up with new ideas if the "idea" appeared as a generalisation durring training or if it was achieved using reason in chain of thought.
Ordinarily in mathematics there’s a TON of papers to write just combining low level problems with different techniques. Better still, and often considered groundbreaking is borrowing techniques from other fields and adapting them or creating analogous methods to solve problems. A lot of landmark papers have been written this way. This is also what transformers are sort of good at within other contexts. They have super human breadth so I’m hopeful they’ll become real assets in math for a long time. Though the leaps necessary to adapt a technique in a nonobvious way might be too much for a while longer. We’ll see.
Truly novel techniques are quite rate indeed and I don’t know if LLMs can represent them faithfully in their embedding space or not. My inclination is that they probably can most of the time, but I don’t know. Mathematicians would describe such thins as “alien”.
A jump in intuition comes from automatic processes reorganising the relational structure of conceptual models. There is no reorganisation of the model durring inference.
one could argue that model can reorganize / interact with prior knowledge captured in text form (edit files) hence there can be reorganisation