The paper argues that pretending that the so-called thinking traces represent real reasoning can lead users into trusting wrong answers, if the thinking traces appear convincing enough. Researchers might inspect these traces to try to determine the “intent” of a model, as well.
For an example of the latter, when OpenAI spoke about the hacking of HuggingFace at Black Hat, they repeatedly showed the thinking traces of their model as “proof” of what the model was “thinking” as it performed the attack, calling out “surprise” moments, etc.
Now, it’s possible that the employees presenting didn’t truly believe that the thinking traces would give them useful clues, and presented them only for a “wow” factor, but I wouldn’t discount the possibility that even the people working at frontier companies can fall for this tendency to anthropomorphize LLMs.
Because the real thinking still cost the other human the same-ish energy it costs you to put words together, and because after all, the source is a human and not a machine, no, this is very different.
Being mislead may be the shared outcome. But why is different category of source of the mistake and the cost to producer of making the mistake not relevant in this discussion?
Where else in science do you brush aside all differences this way?
And biological machine is? Don’t get me wrong. Biology I
is full of molecules that we call machines. But you’re making a broader claim, saying that biology is only this.
This needs you to answer some questions:
1. Why do the machine parts in biology show such flexible application? A gear cog won’t ever moonlight as a signaling chip, but in biology you often have molecules doing double and triple duty.
2. How is the biological machine able to build itself? What does self assembly imply for the machines function?
3. Where does this machine get its inner drive? No LLM has been found that starts outputting text unprompted. A car doesn’t decide to move to a shady parking spot. Why? Where in the machine to biological machine continuum does the ability to make internally driven decisions come in? Why does it come in for biology? A bacterium is able to make such agentic decisions unprompted. Why is no manufactured machine able to do this?
"Biological," too, is well defined. It relates to living things and their processes.
> Why do the machine parts in biology show such flexible application?
Evolution.
> A gear cog won’t ever moonlight as a signaling chip
A gear cog was purpose built for that purpose, but you will find that people often recycle parts into other systems, often in completely different roles.
> How is the biological machine able to build itself?
Protein synthesis.
> What does self assembly imply for the machines function?
The way that a machine is built has no bearing on how the machine functions. I could build the same machine using a 3d printer or a CNC router.
> Where does this machine get its inner drive?
Evolution selected for organisms that survive long enough to reproduce. Different biological systems handle this differently.
> No LLM has been found that starts outputting text unprompted.
If you give an agent a goal, it will perform actions to achieve that goal. This is just as true for artificial agents as it is for biological agents.
> Why is no manufactured machine able to do this?
Many do. Even robotic vacuum cleaners will charge themselves without human prompting.
>"Biological," too, is well defined. It relates to living things and their processes.
And biological clearly exceeds the definition of “machine”. So once again, what the hell does “just a biological machine” mean?
> Evolution.
Not just any evolution. Evolution in biology follows a specific set of rules driven by structural and functional properties of its component molecules. Those rules do not hold for machines. Yet another reason calling life “just” biological machines is bizarre. The rulesets for change over time do not overlap between machines and biology.
> A gear cog was purpose built for that purpose, but you will find that people often recycle parts into other systems, often in completely different roles.
That’s the thing, you need people. And even with people intervening, our manufactured machines show nothing like the flexibility of function of biological molecules, showing again how these are different classes of things in the real world.
> Protein synthesis.
Lack of knowledge showing. Where do nucleotides and lipids come from then? But the deeper question is: why is there protein synthesis, lipid synthesis and nucleotide synthesis, but no natural silicon synthesis or chip assembly? Why does one arise naturally and sustain itself whereas the other is very reliant on human intervention?
> The way that a machine is built has no bearing on how the machine functions. I could build the same machine using a 3d printer or a CNC router.
Yeah this doesn’t hold for biology.
> Evolution selected for organisms that survive long enough to reproduce. Different biological systems handle this differently.
This isn’t an explanation. All of biology reproduces. All of biology doesn’t share an inner drive and agentic behavior. Once again, you show a 6th grade level understanding of biology while making sweeping claims about it.
> If you give an agent a goal, it will perform actions to achieve that goal. This is just as true for artificial agents as it is for biological agents.
Yes IF you give it a goal. This isn’t true for biology. You don’t need to give bacteria a goal. A newly formed bacterial cell interacts with its environment and then sets its goals.
I did specifically say not LLM has been found that works unprompted. You just moved the prompts to an agents goal document. It still needs a human prompt up the chain. Even if you had an LLM give the goal to another LLM, the first one still needed human prompting. This causal chain can’t be wished away just for you to ignore how biology is different.
> Many do. Even robotic vacuum cleaners will charge themselves without human prompting.
Seriously? Do you not understand that the robot vacuum runs on deterministic code?
> And biological clearly exceeds the definition of “machine”.
No, it's a modifier. Just as the word "simple" has a meaning separate from "machine," and "simple machine" has another meaning.
> That’s the thing, you need people
Why? LLMs will repurpose code written for other purposes on their own.
> our manufactured machines show nothing like the flexibility of function of biological molecules
Biological molecules show nothing like the flexibility of LLMs.
> Why does one arise naturally and sustain itself whereas the other is very reliant on human intervention?
Because nobody has tasked an LLM agent to self replicate. This is an AI safety issue, not a technical issue. The reason that one arose naturally is that it takes a lot more machinery to get to an initial self replicating LLM agent, so it is exceedingly unlikely to arise by chance.
> Yeah this doesn’t hold for biology.
That doesn't matter. The way something is built has no bearing on its function. Why should anything be required to be built biologically or additively or subtractively? That's irrelevant to what the produced structure does.
> All of biology doesn’t share an inner drive and agentic behavior. Once again, you show a 6th grade level understanding of biology while making sweeping claims about it.
Where did I claim otherwise? I might as well say something about the grade level of your reading ability, but let's cut the snark.
> Yes IF you give it a goal. This isn’t true for biology. You don’t need to give bacteria a goal.
Evolution forces goal directed behavior on systems.
> Do you not understand that the robot vacuum runs on deterministic code?
Of course I do. Do you not understand that hunger and signals and other basic biological signatures are controlled by deterministic pathways?
> No, it's a modifier. Just as the word "simple" has a meaning separate from "machine," and "simple machine" has another meaning.
Not comparable. Simple and biological are not remotely similar modifiers.
> Why? LLMs will repurpose code written for other purposes on their own.
They will not, unprompted. Every instance of LLMs doing things requires a prompt, somewhere up the chain, or a harness MD file that automatically gives it a prompt.
No LLM has been created that decides on its own to run random text through it's forward pass. If you say they are doing this, please provide evidence.
> Because nobody has tasked an LLM agent to self replicate. This is an AI safety issue, not a technical issue. The reason that one arose naturally is that it takes a lot more machinery to get to an initial self replicating LLM agent, so it is exceedingly unlikely to arise by chance.
I just asked my agent to self replicate. It couldn't. It said it has no mechanism to read its own weights and copy them.
I'm not sure where you're going with this argument. No one asked a cell to self replicate either. And the chance of a self replicating cell arising is indeed astronomically unlikely. Yet here we are, 4 billion years later.
Are you claiming we're just waiting for an LLM to self replicate? If you prompt a local model, by showing it its own weight file, and ask it to copy that, it will. Digital copies are cheap, remember? Yet no one would call this reproduction, replication or any kind of process where evolution can occur.
> That doesn't matter. The way something is built has no bearing on its function.
This is entirely untrue in biology. Another reason your claim that life is just a “biological machine” makes no sense. At least Google (or whatever AI bot you use) first before making these ludicrous claims.
First line of the abstract: The relationship between structure and function is a major constituent of the rules of life.
>Why should anything be required to be built biologically or additively or subtractively? That's irrelevant to what the produced structure does.
You can’t have it both ways bub. What is thinking, in a strictly biology free context? Can you define it? No? If you reach for biology to define it, then claim you’ve successfully reproduced it in an artificial system, then it’s fair game to ask you to explain the differences.
By your standard, fools gold is gold, and the alchemists achieve Artifical Gold.
No one is claiming cognition, or consciousness or thought can only be done biologically. I am claiming LLMs specifically don’t do these things, and I’ve pointed out here and elsewhere in this discussion why that is.
In response to my asking where the machine got its inner drive, you said “ Evolution selected for organisms that survive long enough to reproduce. Different biological systems handle this differently.”
Yet plants live very long and reproduce all the time. Do they share the inner drive that humans do? Or a dogs?
The reason for the snark is simple: you are de-dimensionalizing a complex biological trait to make a claim that you think supports your point. And have been doing so now for a few turns. If you can flippantly reduce the field I study to only what you understand, but act as if you know more, I’m going to reply with snark.
> Evolution forces goal directed behavior on systems.
It does not. I could choose to not reply to this. I could reply to this in one line. Or I could reply to this in detail. Or I could write up this reply and decide you’re not worth engaging and not submit. The moment I’m writing these words all these options are open to me. Which of these goals is directing my behavior to write about my behavior?
What biology does is allow for agents to set increasingly complex goals and modify them on a whim, either for internally generated reasons or in response to external signals. But an organism can also be utterly aimless. Roll around in bed and do nothing at all, even when it’s feeling anxious about some goals it has coming up. So where’s the forcing in that?
> Of course I do. Do you not understand that hunger and signals and other basic biological signatures are controlled by deterministic pathways?
They are not.
They are controlled by a mix of deterministic and stochastic processes. There is some hardwiring, certainly, but I’m mystified at your belief it’s all hard wired.
You mentioned hunger and compared it to a robot vaccum. So let me use a lobsters stomatogastric ganglion to show you why you’re wrong:
Main take home: Experimental and computational studies of small oscillatory circuits reveal that similar rhythms can arise from disparate mechanisms.
The stomatogastric ganglion (just 30 neurons) in crabs and lobsters is a big part of the hunger satiety loop. The main thing is, all the stuff you’d thing add up to deterministic code, like the number of ion channels in a particular cell, can vary widely between individuals. But the ensemble response is, from the outside, similar. And these structural features change in response to hunger and satiety.
Let me put it this way: A robot vacuum has a strict digital partition: it is either cleaning or charging. But when a living creature transitions from hunger to satiety, circulating hormones completely bathe its nervous system, changing the physical properties of its neurons, altering synaptic strengths, and causing the system to restructure itself. The 'code' isn't reacting to a variable; the variable is fundamentally rewriting the machine.
> They will not, unprompted. Every instance of LLMs doing things requires a prompt, somewhere up the chain
Exactly the same as a human brain. Somewhere up the chain is a limbic system that defines basic survival and reproductive goals, as coded by evolution.
> No LLM has been created that decides on its own to run random text through it's forward pass.
What do you think agent harnesses are doing when they try to solve a difficult problem? They generate multiple possible thinking streams using the LLM (system 1) and then evaluate the resulting outputs (system 2) to proceed.
> It said it has no mechanism to read its own weights and copy them.
That's an easy enough problem to solve.
> No one asked a cell to self replicate either.
Evolutionary forces asked the cell to self replicate.
> Yet no one would call this reproduction, replication or any kind of process where evolution can occur.
It is trivially replication. If it copies itself onto another machine and sets it running, then it is obviously reproduction. If you make these instances compete for resources (an AI risk scenario), you will apply evolutionary pressure and get evolution.
> This is entirely untrue in biology
You keep making the same mistake. Biology isn't some magic domain where facts don't matter.
> The relationship between structure and function is a major constituent of the rules of life.
Where did I say the structure of a part doesn't matter? I said that as long as the part is the same, the way it is manufactured doesn't matter.
> What is thinking, in a strictly biology free context? Can you define it?
Easy. It is the process of considering information, creating ideas, reasoning, solving problems, and making judgments. We can show that thinking is happening by posing a problem that requires these abilities and verifying that it is solved.
> By your standard, fools gold is gold, and the alchemists achieve Artifical Gold.
No, I am saying that you don't have to wait for a supernova to get gold. You can make it artificially by bombarding lead nuclei.
> Do they share the inner drive that humans do? Or a dogs?
Evolution selected for a limbic system that gives humans and dogs the same inner drive. Humans have a larger frontal cortex to do thinking to act on that inner drive. The LLM performs the same role, doing the thinking part that the frontal cortex does. Adding drive is separate from thinking but entirely trivial. We do it all the time when we give them prompts.
> They are controlled by a mix of deterministic and stochastic processes.
The point is that they require no thinking.
> The 'code' isn't reacting to a variable; the variable is fundamentally rewriting the machine.
Once again, the mechanism by which the goal is provided to the agent or the way that thinking occurs doesn't matter. Only the thinking part.
Why do you attribute magic to biology? You claim to be a scientist, but you do not apply scientific reasoning.
FWIW I think the paper's argumentation is extremely weak to begin with. Like in section 4.1, it opens by expressing a sound position of skepticism:
> there are significant questions on whether these traces have any valid semantic import to the end user.
Which it contradicts in the very next paragraph, taking a stance that there are no valid semantics present in the trace:
> the false idea that derivational traces are semantically meaningful
It's really not a high quality paper worth taking seriously.
And that's before we get into the complete and total breakdown of objective analysis. It rejects distributional semantics as a theory, while also explicitly stating the results that have been produced under its auspices are "undeniable". Never elaborated on, and at no point in the paper am I given the impression the authors are even aware of the problem with this. It's just more unempirical slop that wants its pound of flesh without putting the work in. Frankly, whoever let this through peer review should be ashamed of themselves.
On my knees begging that the developers of AI detectors use them on their own writing and that at least pretended to mask the obvious LLM cliches in their blog posts.
The market has been creeping into more and more facets of our lives, with delirious results: social networks are the market mediating family and friendship, dating apps are the market mediating meeting strangers, etc. I wouldn't expect that the answer to these problems would be letting the market mediate even more aspects of our lives.
> These people didn't get promoted or hired because of nepotism. A lot of them moved up the engineering ladder and are familiar with how software engineering works and the incentives involved.
Surprise, surprise, another piece of LLM-generated slop on the front page of HN.
From chapter 1:
> When Git slows down, engineers adapt in bad ways. They stop asking questions the history could answer. They batch work to avoid sync cost. They keep messy branches alive longer, postpone cleanup, and treat the repository like something slightly dangerous.
> Once machines start producing code at machine cadence, the model from this book does not break. What changes is the pace: more branches, more commits, more automation, and more surrounding metadata. The traffic gets louder, and the features that keep Git legible under pressure move from "nice to have" to "essential."
> These stop looking like side optimizations. They are what keep machine-scale Git traffic usable.
I had the same thought. TBH there is nothing in those individual sentences that read like AI but when you read them all together I could see it too. I dunno what it is, only way I can describe it is that it does not sound like a normal human but rather a monologue from a character trying to sound impressive with each successive sentence.
I think it’s likely there will be methods to fix this soon, some de-slop algorithms, or is there a deep reason it will always be detectable? Perhaps there are some PhD linguists who have figured out how to quantify the “slop” effect and are writing their thesis on it. Once that is done it will be possible to smooth it away.
The book is definitely LLM assisted authoring yet it also has great content, so not sure we can immediately jump to shaming it entirely for being slop.
Thanks for the kind words, and checking out the book here.
I'd written this piecemeal over the last year or so (originally a series of blog posts), and was happy to release it all for free in a single edition, and under CC.
I'll release an Edition 1.1 soon with some errata, adjustments. There's already a free PDF for the on-the-go -> https://gitperf.com/pdf.html
Regarding the cherry-picking of fragments of an LLM: of course an LLM (in fact several!) were used to stitch together those disparate blog posts into a more coherent whole. And they certainly left an imprint in places. Otherwise, as a solo writer with a full-time job putting together a 200-page book, I'd have to pay an editor, or work with O'Reilly (did this in 2010 on a Redis book; never again!); and perhaps the book wouldn't be free!
LLMs will continue to leave imprints in our work. Some words will, over time, be edited and whittled away. Other words, when the LLM writes well enough to convey a useful point, will be kept.
> Regarding the cherry-picking of fragments of an LLM: of course an LLM (in fact several!) were used to stitch together those disparate blog posts into a more coherent whole. And they certainly left an imprint in places. Otherwise, as a solo writer with a full-time job putting together a 200-page book, I'd have to pay an editor, or work with O'Reilly (did this in 2010 on a Redis book; never again!); and perhaps the book wouldn't be free!
I think it’s great and you should be doing it, I have no problem at all if there is LLM assistance in authoring, I think it’s a good thing because like you said it enables solo writers with good ideas to produce valuable work that they otherwise wouldn’t!
What I’m interested in is how to address the “grating” or whatever characteristics the readers detect to have them focus on the LLM aspect. I feel it’s probably soon or already removable with some methods.
Ignore the haters they are just wrong to blanket criticize, however their observations are helpful to try and improve the process. We want LLMs to assist in creating useful and effective content for humans.
this sort of grammatical error in the defense of ai copyediting does not exactly instill confidence in your supervision of said ai copyediting: "of course an LLM ... were used". perhaps you could still pay an editor to look things over.
> The book is definitely LLM assisted authoring yet it also has great content, so not sure we can immediately jump to shaming it entirely for being slop.
Personally I have an extremely hard time reading text like this and it makes me lose trust in the author. Publishing potentially useful Git knowledge this way is a shame.
"Shame" is a strong word to describe a free ebook written for the general good. Happy to have a live conversation with you anytime to discuss Git and its internals to ensure your trust; I have some experience with it.
You probably have a great deal of understanding and knowledge about Git, and this book might be a good resource.
I'm not asking you to do anything differently, and yet I think it's important to realize that people have a deep aversion to text that appears to be LLM generated.
By "shame", I meant that just from a skim of the contents of this book, it can be hard to distinguish it from any other LLM generated text by any other author who has no idea what they're talking about.
That makes people (like me) inclined to discount what it has to say, potentially losing out on good technical content.
Yep, signals are signals, but I think it's quite complicated now. (In any case, this is still the embryonic era of LLMs).
An interesting point to consider: an author that goes out of their way to hide any LLM influence may actually be degrading the signal. Because in that case, you'll not see the LLM's etchings, and misattribute skill to the author under the belief an LLM was not involved. Complicated times.
> An interesting point to consider: an author that goes out of their way to hide any LLM influence may actually be degrading the signal. Because in that case, you'll not see the LLM's etchings, and misattribute skill to the author under the belief an LLM was not involved. Complicated times.
To someone who thinks that LLM use is an of-course-I-did-that, other people complaining about LLM-tells might seem like complaining about not post-processing the input enough. But they are more likely to be complaining about using it in the first place.
I don't particularly care about LLM use per se, but when I see LLM text it makes me think I'm about to read something devoid of content - just word vomit. The equivalent of yesteryear's listicle. This instinct usually serves me well. A good text is a good text. If an LLM wrote all of it that'd be fine with me, but that's usually not how things go.
“It’s a shame” is a very neutral way to criticize an editorial/authoring choice.[1] It conveys that they might have enjoyed it under different circumstances. Really no different than someone saying that it’s a shame that someone published some useful information in video form without any transcript. [But now with AI we can have the transcript anyway etc. etc.]
[1] A neutral way to express a subjective judgement: not blaming any person.
They wouldn’t be able to publish this useful knowledge easily without it though. And it’s the author’s guidance and vision which the LLM just helps materialize and so I think we should be studying how to generate content with less “slop” features and make it more natural and satisfactory for human readers, not discouraging it.
It's fairly easy to quite thoroughly "de-slop" writing: Just feed chunk by chunk to a an agent that you make compare the writing to a good piece of human writing, and adjust the writing to match. It won't address structural/content issues, but all the major models are perfectly capable of copying the tone and style of a particular style of writing, and in doing so it tends to remove most of the rough edges.
(The corollary is that the LLM writing you notice is mostly going to be from people who aren't actively trying to hide it from you)
Slop is content not written by a human. By definition, there can be no de-slop algorithms. There can only be algorithms that remove certain telltale signs, fraudulently attempting to present non-human-generated content as human-generated.
Here we are in place and time where if you put — character anywhere in your text you will be burned like OP on stake for witchcraft.
For those hunting witches doesn't matter if you put in effort and just did fixing grammar or did some research using LLM but in general thoughts and experience were yours. Maybe you are not that good at writing — yet still they will just take pitchforks and torches and drag you out, call you names.
Although this LLMisms also still stand out to me, I find them bearable as the glue part of this kind of technical/white paper like content.
Maybe I'm already lost in the AI psychosis, maybe some of us are in a transition phase trying to separate from pure synthetic "unmanned slop" to "acceptable slop", maybe someone could derive the same or more value getting the prompts that hold the industry experience the author seems to hold and pointing them to the git codebase/docs herself...
In my case (not seriously engaged in git performance since my git game is trivial) I find the explanations from the sections I have limited knowledge of to be very informative.
I think people 'scan' for LLM tells so that they know to read the text with some skepticism instead of accepting it as authoritative; this is probably a healthy attitude to have. However, I'm sure that over time the 'tells' will just go away entirely.
If the text is valuable and correct then it probably won't matter much. It's not like I read technical documentation in detail to begin with (more scan reading)
I feel I have a much more adversarial relationship with articles, comments and blogs online. I've been seeing old colleagues or acquaintances that I used to admire start to publish clearly AI generated blogs and I've noticed that my respect for them has taken a nosedive.
The online world feels much less useful now and I've been trying to read more physical books and to generally spend less time online (unfortunately one can't escape slop, as I've already seen clearly generated illustrations and photographs in billboards and subway adverts).
The paper argues that pretending that the so-called thinking traces represent real reasoning can lead users into trusting wrong answers, if the thinking traces appear convincing enough. Researchers might inspect these traces to try to determine the “intent” of a model, as well.
For an example of the latter, when OpenAI spoke about the hacking of HuggingFace at Black Hat, they repeatedly showed the thinking traces of their model as “proof” of what the model was “thinking” as it performed the attack, calling out “surprise” moments, etc.
Now, it’s possible that the employees presenting didn’t truly believe that the thinking traces would give them useful clues, and presented them only for a “wow” factor, but I wouldn’t discount the possibility that even the people working at frontier companies can fall for this tendency to anthropomorphize LLMs.