Rendered at 18:57:58 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
StilesCrisis 21 hours ago [-]
Had to stop reading when the article devolved into Claude spam. "defensible in isolation," "honestly ranked," ugh. Please write your own blog post.
robwwilliams 19 hours ago [-]
Yep. Such a classic Claude flavored product. Good ideas wrapped in cotton candy.
I often read as much as 1000 words thinking to myself: “this is smooth”, but then think to myself “Is this my friend Opus 4.8 or now Opus 5?.
In this case the rhetorical neatness is unmistakable especially in the beginnings and endings of paragraphs: “Here is the constructive turn, honestly ranked, with no silver bullets on offer.”
Yes: and that is actually a smoking gun.
edavison1 18 hours ago [-]
This is the second ACM article in the last few months that was clearly not written by a human. Funny because it's against their editorial policy... what is going on over there?
renlo 20 hours ago [-]
My strategy these days is to scan and look for the tells and click out when I see them. Mine was the same "honestly ranked, with no silver bullets on offer". I suspect in less than a year we won't be able to tell the difference.
rogerrogerr 20 hours ago [-]
I’ve thought this for a while, but why hasn’t it happened yet? At this point, OpenAI and Anthropic and friends could definitely remove the AI “smell” from writing output, or give users a first class way to specify a writing style.
So why haven’t they? My theory is they see this as a sort of fingerprint, useful to not train on later. Or something. Maybe they just don’t care. Certainly feels either intentional or a result of ambivalence.
It’s certainly true today that I probably wouldn’t know an AI written article if the author went out of their way to use one of the many prompts available to tone down the AI-isms.
This wiki page is updated frequently and only needs to be fed into an AI to remove its telltale writing.
xboxnolifes 18 hours ago [-]
If the underlying prompt of the model stays the same, it seems to me that LLMs will always have common tells unless overridden with a thorough prompt from the end user. It's like if you had 1 person write half of the content on the internet. You'd probably get pretty good at noticing their writing style.
Maybe my understanding of LLMs is wrong, but it seems obvious to me that when you have a large corpus of LLM output you will eventually notice common tells when everyone is using the same models, weights, and base prompt.
thatjoeoverthr 19 hours ago [-]
It’s natural to the fact it’s the same model. Everyone has tics, and when a given model is asked to write millions of texts, they become visible. But! Some portion of the audience and user base can’t see it, so there is no benefit to fixing it. Case in point, yet another hustler felt very clever posting slop, and the likely actual audience (Google’s ranking system) probably does like it.
paulpauper 19 hours ago [-]
Because it's doing what is does best: making predictions, which works great when you're not being judged on aesthetics, such as coding or math, but the human thought process is messy or erratic. It just falls apart when you do the next token process to it. If the goal is to "convey information in readable chunks," AI does great at this.
robwwilliams 19 hours ago [-]
I have started beating the hell out of Claude with stylometric analyses of writers I admire; usually technical writers like Terry Winograd, Leslie Lamport, Rodney Brooks. Then the “no or minimal rhetorical flourishes” rule; no British em-dashes, and minimize the negative phrase thesis-antithesis fun”. It helps.
I think you are right that in a year or two LLMs will be able to do a good impression of many technical styles. But not Nabokov, Kundera, or Kafka for subtlety.
If one manages to channel Edsger W. Dijkstra I will be impressed and rank it high on my leaderboard.
frollogaston 17 hours ago [-]
This article was too verbose for me to want to read it, but I'm still not sure about it being LLM-gen'd. All I got is Claude says "honest" a lot.
cephei 20 hours ago [-]
In a year or more, the audience will likely be an agent instead of a person. It'll be interesting to see how that shifts language and article formats.
thatjoeoverthr 19 hours ago [-]
It has been for years. The actual audience for a great deal of text you see is Google’s ranking system. Look up any recipe and ask yourself who reads the ten paragraph story time about grandma’s cookies. It literally isn’t intended to be read.
StilesCrisis 19 hours ago [-]
I believe this also partly sprung up because recipes in isolation don't qualify for copyright. The flavor text gives you grounds to sue if a clone of your recipe site pops up somewhere else.
paulpauper 20 hours ago [-]
It already is. AI bots overflowing visitor logs now with endless IPs
paulpauper 20 hours ago [-]
lol
reminds me of a Claude math paper:
"honestly sharp , no hype: cos(pi+pi)+2+2=cos(2pi)+4=1+4=5"
paulpauper 20 hours ago [-]
Yeah, the most obvious giveaway is to ask yourself, "is this how a human would actually write?"
hexagonwin 10 hours ago [-]
lmao, even GPTzero gives 100% on a snippet I copied. (from "Here is the constructive turn" to "and slower is what we can actually buy.)
throw10920 16 hours ago [-]
The obvious solution is to have non-public benchmarks.
It is exceedingly difficult to train on a proprietary benchmark administered by someone with half a brain (i.e. don't sign up for a ChatGPT account with your benchmark@artificialanalysis.ai email) - you have to find a tiny needle in a vast haystack.
In fact, it can be difficult enough that it's simply not economically viable - that is, that it's cheaper to make the model better than it is to try to find the account running the benchmark.
In the limit case, the benchmark is indistinguishable from...normal problems that need to be solved.
astro1234 23 hours ago [-]
I agree and that’s why we need and indeed have an ever evolving landscape of benchmarks
> Private, refreshed test sets attack the mechanism itself, and in my view they are the only intervention that does. If the questions have never touched the public Web, they can’t be in the training data; if they rotate, memorizing this year’s set doesn’t help next year.
That’s what we have. A fresh public benchmark is also good, and teams do make efforts to decontaminate training data but there’s likely just no great way around leakage.
Btw, lots more issues in benchmarks than the ones discussed; for instance you can leak answers from the questions themselves or in the case of e.g. multiple choice formats in the actual answers. You just pass the MCQ choices themselves to the model and it may be able to guess way above chance. Coding agent benchmarks sometimes forget to delete .git. They mention e.g. a 6.9% error rate in one of the benchmark items, this seems pretty typical and I would actually be fine shipping that.
Benchmarks are very ugly, but if they didn’t exist we would need to invent them. All of the problems above and more do not explain the progress we see. There are probably 50,000 benchmarks in the literature and new ones get created frequently with varying levels of quality and usefulness.
zahlman 22 hours ago [-]
Clearly, the solution is to judge society by how many currently-un-gamed benchmarks it has produced.
froh 12 hours ago [-]
this one numbers, man...
lstodd 18 hours ago [-]
Benchmark is by definition gamed. That is the essence of Goodhart's law.
Society is a benchmark of sorts. You see where that leads.
----
edit: as in this quote from Gulliver:
I told him, “that in the kingdom of Tribnia, by the natives called Langdon, where I had sojourned some time in my travels, the bulk of the people consist in a manner wholly of discoverers, witnesses, informers, accusers, prosecutors, evidences, swearers, together with their several subservient and subaltern instruments, all under the colours, the conduct, and the pay of ministers of state, and their deputies. The plots, in that kingdom, are usually the workmanship of those persons who desire to raise their own characters of profound politicians; to restore new vigour to a crazy administration; to stifle or divert general discontents; to fill their coffers with forfeitures; and raise, or sink the opinion of public credit, as either shall best answer their private advantage. It is first agreed and settled among them, what suspected persons shall be accused of a plot; then, effectual care is taken to secure all their letters and papers, and put the owners in chains. These papers are delivered to a set of artists, very dexterous in finding out the mysterious meanings of words, syllables, and letters: for instance, they can discover a close stool, to signify a privy council; a flock of geese, a senate; a lame dog, an invader; the plague, a standing army; a buzzard, a prime minister; the gout, a high priest; a gibbet, a secretary of state; a chamber pot, a committee of grandees; a sieve, a court lady; a broom, a revolution; a mouse-trap, an employment; a bottomless pit, a treasury; a sink, a court; a cap and bells, a favourite; a broken reed, a court of justice; an empty tun, a general; a running sore, the administration.
When this method fails, they have two others more effectual, which the learned among them call acrostics and anagrams. First, they can decipher all initial letters into political meanings. Thus N, shall signify a plot; B, a regiment of horse; L, a fleet at sea; or, secondly, by transposing the letters of the alphabet in any suspected paper, they can lay open the deepest designs of a discontented party. So, for example, if I should say, in a letter to a friend, ‘Our brother Tom has just got the piles,’ a skilful decipherer would discover, that the same letters which compose that sentence, may be analysed into the following words, ‘Resist -, a plot is brought home - The tour.’ And this is the anagrammatic method.”
== some ecclesiastes quote would be nice here.
ShadowOfThePit 13 hours ago [-]
I don't understand your point, or your quote.
Yes, if you try hard enough, you will find proof for any accusation whose outcome you've already decided. But what does that have to do with anything about society being a benchmark?
throwatdem12311 20 hours ago [-]
This reminds me of many years ago when Mozilla/Firefox (I think it was) said that they stopped focusing on mainstream benchmarks because they didn’t really translate to real world browser performance gains.
I view these AI benchmarks the same. No I do not care that GPT got 1200 on FartAGIMaX-4.0-Extreme and Claude got 1350. I care about how much it costs and how correctly it does the tasks that I give it. Unfortunately the only way to know is to use them all myself and measure it myself.
At the end of the day these things are all so damn close in how they behave in whatever harness so it realy just does boil down to whatever is actually cheapest.
This is why Deepseek is great: it’s so much cheaper it doesn’t matter if I burn way more tokens because it’s still orders of magnitude cheaper than the US SotA models. If it doesn’t get it quite right immediately I just do a few more turns and then it’s fine. Barely an inconvenience.
ddp26 21 hours ago [-]
Not forecasting though. You can't goodhart predicting real-world events
teddyh 23 hours ago [-]
“When you place a tangible value on trust, trust becomes a commodity to be bought and sold.”
I think Goodhart's law is just a consequence of correlation vs causation.
It is very easy to find a metric that is correlated with what you want. But once you start trying to influence a system, you quickly push it out of the range where the correlation holds.
In order to optimize for something, you need to maximize the actual causative variable. This is much harder.
abdullahkhalids 22 hours ago [-]
That is correct. It's just that in most cases in the real world, there is a complex casual network, and many of those variables are not even measurable. So you have to pick a proxy for one or more of the variables and make a metric out of this.
The other problem is that in the real world, we want to make decisions, and the easiest way to make decisions is to have a single metric to judge everything by. With multiple metrics, you get into these debates about subjectivity.
You can get around Goodhart's law if you are able to pick multiple proxy variables and demand that the user optimize them all. And you pick these variables in a way that it's really hard to cheat (i.e. deoptimize the actual intended variable while optimizing the proxy variables). Game designers do this all the time for example, because the system is clean and simple enough to do it.
jldugger 23 hours ago [-]
I think the difference is that Goodhart's law describes how the causal chain _changes_ as a result of management behavior, and in particular the incentives they design for the labor they manage. Incentives are a causal variable for outcomes, and what happens is that people find much easier ways to produce the outcomes you thought you wanted.
Like if you manage a call center and set up KPIs around average call time, reps will start hanging up on customers. Employees could always have done that, and the causal link was always there, there was just no reason to.
IMO the problem is executives want (and perhaps need) their directs to report and track one big number month over month. If you give them five metrics they'll never know if you're making progress or just oscillating between a few local minima. And if each of their ten directs has five metrics, you now have 50 numbers and no idea what time it is[1].
Everything, every. single. thing. that can cause people to spend money on, will be ultimately manipulated.
Will be true as long as humanity exists.
cyanydeez 22 hours ago [-]
obviously, the best benchmark is the one you tell no one about.
functionmouse 23 hours ago [-]
Jokes on them, I don't trust benchmarks
Once something becomes a benchmark it is no longer a good benchmark.
fluoridation 23 hours ago [-]
No, it's when a benchmark becomes a target. You might have a private benchmark that you tell no one about. Would you not trust it?
charlieyu1 19 hours ago [-]
Good benchmarks are costly to build even for mid-large corporations. And once the benchmark is used on models you really can’t tell if the problems would be scrapped for training
I often read as much as 1000 words thinking to myself: “this is smooth”, but then think to myself “Is this my friend Opus 4.8 or now Opus 5?.
In this case the rhetorical neatness is unmistakable especially in the beginnings and endings of paragraphs: “Here is the constructive turn, honestly ranked, with no silver bullets on offer.”
Yes: and that is actually a smoking gun.
So why haven’t they? My theory is they see this as a sort of fingerprint, useful to not train on later. Or something. Maybe they just don’t care. Certainly feels either intentional or a result of ambivalence.
It’s certainly true today that I probably wouldn’t know an AI written article if the author went out of their way to use one of the many prompts available to tone down the AI-isms.
This wiki page is updated frequently and only needs to be fed into an AI to remove its telltale writing.
Maybe my understanding of LLMs is wrong, but it seems obvious to me that when you have a large corpus of LLM output you will eventually notice common tells when everyone is using the same models, weights, and base prompt.
I think you are right that in a year or two LLMs will be able to do a good impression of many technical styles. But not Nabokov, Kundera, or Kafka for subtlety.
If one manages to channel Edsger W. Dijkstra I will be impressed and rank it high on my leaderboard.
reminds me of a Claude math paper:
"honestly sharp , no hype: cos(pi+pi)+2+2=cos(2pi)+4=1+4=5"
It is exceedingly difficult to train on a proprietary benchmark administered by someone with half a brain (i.e. don't sign up for a ChatGPT account with your benchmark@artificialanalysis.ai email) - you have to find a tiny needle in a vast haystack.
In fact, it can be difficult enough that it's simply not economically viable - that is, that it's cheaper to make the model better than it is to try to find the account running the benchmark.
In the limit case, the benchmark is indistinguishable from...normal problems that need to be solved.
> Private, refreshed test sets attack the mechanism itself, and in my view they are the only intervention that does. If the questions have never touched the public Web, they can’t be in the training data; if they rotate, memorizing this year’s set doesn’t help next year.
That’s what we have. A fresh public benchmark is also good, and teams do make efforts to decontaminate training data but there’s likely just no great way around leakage.
Btw, lots more issues in benchmarks than the ones discussed; for instance you can leak answers from the questions themselves or in the case of e.g. multiple choice formats in the actual answers. You just pass the MCQ choices themselves to the model and it may be able to guess way above chance. Coding agent benchmarks sometimes forget to delete .git. They mention e.g. a 6.9% error rate in one of the benchmark items, this seems pretty typical and I would actually be fine shipping that.
Benchmarks are very ugly, but if they didn’t exist we would need to invent them. All of the problems above and more do not explain the progress we see. There are probably 50,000 benchmarks in the literature and new ones get created frequently with varying levels of quality and usefulness.
Society is a benchmark of sorts. You see where that leads.
---- edit: as in this quote from Gulliver:
I told him, “that in the kingdom of Tribnia, by the natives called Langdon, where I had sojourned some time in my travels, the bulk of the people consist in a manner wholly of discoverers, witnesses, informers, accusers, prosecutors, evidences, swearers, together with their several subservient and subaltern instruments, all under the colours, the conduct, and the pay of ministers of state, and their deputies. The plots, in that kingdom, are usually the workmanship of those persons who desire to raise their own characters of profound politicians; to restore new vigour to a crazy administration; to stifle or divert general discontents; to fill their coffers with forfeitures; and raise, or sink the opinion of public credit, as either shall best answer their private advantage. It is first agreed and settled among them, what suspected persons shall be accused of a plot; then, effectual care is taken to secure all their letters and papers, and put the owners in chains. These papers are delivered to a set of artists, very dexterous in finding out the mysterious meanings of words, syllables, and letters: for instance, they can discover a close stool, to signify a privy council; a flock of geese, a senate; a lame dog, an invader; the plague, a standing army; a buzzard, a prime minister; the gout, a high priest; a gibbet, a secretary of state; a chamber pot, a committee of grandees; a sieve, a court lady; a broom, a revolution; a mouse-trap, an employment; a bottomless pit, a treasury; a sink, a court; a cap and bells, a favourite; a broken reed, a court of justice; an empty tun, a general; a running sore, the administration.
When this method fails, they have two others more effectual, which the learned among them call acrostics and anagrams. First, they can decipher all initial letters into political meanings. Thus N, shall signify a plot; B, a regiment of horse; L, a fleet at sea; or, secondly, by transposing the letters of the alphabet in any suspected paper, they can lay open the deepest designs of a discontented party. So, for example, if I should say, in a letter to a friend, ‘Our brother Tom has just got the piles,’ a skilful decipherer would discover, that the same letters which compose that sentence, may be analysed into the following words, ‘Resist -, a plot is brought home - The tour.’ And this is the anagrammatic method.”
== some ecclesiastes quote would be nice here.
Yes, if you try hard enough, you will find proof for any accusation whose outcome you've already decided. But what does that have to do with anything about society being a benchmark?
I view these AI benchmarks the same. No I do not care that GPT got 1200 on FartAGIMaX-4.0-Extreme and Claude got 1350. I care about how much it costs and how correctly it does the tasks that I give it. Unfortunately the only way to know is to use them all myself and measure it myself.
At the end of the day these things are all so damn close in how they behave in whatever harness so it realy just does boil down to whatever is actually cheapest.
This is why Deepseek is great: it’s so much cheaper it doesn’t matter if I burn way more tokens because it’s still orders of magnitude cheaper than the US SotA models. If it doesn’t get it quite right immediately I just do a few more turns and then it’s fine. Barely an inconvenience.
— <https://news.ycombinator.com/item?id=27432186>
It is very easy to find a metric that is correlated with what you want. But once you start trying to influence a system, you quickly push it out of the range where the correlation holds.
In order to optimize for something, you need to maximize the actual causative variable. This is much harder.
The other problem is that in the real world, we want to make decisions, and the easiest way to make decisions is to have a single metric to judge everything by. With multiple metrics, you get into these debates about subjectivity.
You can get around Goodhart's law if you are able to pick multiple proxy variables and demand that the user optimize them all. And you pick these variables in a way that it's really hard to cheat (i.e. deoptimize the actual intended variable while optimizing the proxy variables). Game designers do this all the time for example, because the system is clean and simple enough to do it.
Like if you manage a call center and set up KPIs around average call time, reps will start hanging up on customers. Employees could always have done that, and the causal link was always there, there was just no reason to.
IMO the problem is executives want (and perhaps need) their directs to report and track one big number month over month. If you give them five metrics they'll never know if you're making progress or just oscillating between a few local minima. And if each of their ten directs has five metrics, you now have 50 numbers and no idea what time it is[1].
[1]: https://en.wikipedia.org/wiki/Segal%27s_law "A man with two watches never knows what time it is"
Will be true as long as humanity exists.
Once something becomes a benchmark it is no longer a good benchmark.