One of the things I increasingly wonder about is how many of the characteristics and personas of the various models are due to the underlying model itself (ie the initial weights) vs the various post training strategies that are used (ie someone steered the model in one way or another). Or similarly how many of the clear limitations of AI writing are solvable with different post training strategies. Basically are these things fundamental to the AI models themselves, or some layer on top of it that has made the models think this is right. It’s interesting that some things are consistent across model providers and other things aren’t, eg ChatGPT and Claude have very different voices. With time as more models come out, esp open weight ones, I suspect there will be more variation in the post training that allows us to see. I suspect any domain that’s subjective and “taste” heavy will get disproportionate weight from post training, which makes the role of those people doing the tuning especially impactful. And of course as more of the training data becomes AI generated, then the two become harder to distinguish
Uh oh ... did you redact the em-dashes? Or use them as a red herring? It's like the scene from Princess Bride. "He deliberately used em-dashes in Marmoset to fool us into thinking Rabbit was human. But then he knew we were going to assume this so he did the opposite and guided the LLM to liberally use m-dashes. But of course he knew we would assume he was trying to deceive us by thinking he was deceiving us when he wasn't ...."
"Never go in against a Sicilian when death is on the line!"
Uncertainty is a lovely human quality, one of many. Unequivocal certainty, on the other hand, is hubris or ignorance or both. AI, in my limited one-person experience, professes certainty until you point out something it has written that's provably wrong, and then it's politely apologetic. AI isn't embarrassed about its mistakes the way I would be. It hasn't had my awkward adolescence or the striving for excellence instilled in me by my parents, or the trial-and-error learning that has given me opinions about how I want to sound now, which is quite different from how I would have wanted to sound twenty years ago. Uncertainty is but one of the qualities that comes with a beating human heart.
LLMs owe everything they "know" to the collective writings of a vast array of humans (without giving those humans credit or compensation, by the way.) We humans may be intellectual and logical (at our best) but we are also deeply and unavoidably emotional. LLMs, like vampires, suck the blood from the writing of living humans and their writing sometimes appears to be alive, but it's not alive in the way of something human-written. (Don't get me wrong; LLMs are often very useful vampires.) With their encyclopedic access to info and improving grasp of syntax and rules of argumentation, LLMs can be immensely helpful in shaping and refining what we humans write but ultimately, as Claude admitted to me, "I don't have a heart for things to land in."
Our human language is a unique adaptation in the animal world, but it's our animal nature that makes good writing exhilarating. Reading and writing, at their best, are profound emotional experiences; exhilaration is an emotion! Okay, maybe AI can generate a short story that wins an award, even makes you smile like "sunrise over a sink" (wtf?) or maybe human judges are sometimes wrong, and their decisions falter like sunset over a kitchen sink disposal. Mr. Maynard, I'd trust your writing, with or without uncertainty, over LLM output, regardless of how attractive the undead may sometimes be.
It's a well known bias of LLMs that they prefer their own outputs to those of other models and to those of humans: https://arxiv.org/pdf/2404.13076. so I wouldn't put too much credence in an LLM telling you that a paper written by an LLM is better to a human written paper.
Thanks - hadn’t seen this paper. Of course this is a moving target with the sophistication of agentic frontier AI, but clearly from my experience still an issue. That said there is a rapidly growing community of people who take LLM critique as the gold standard …
Nice post. I’ve used AI assisted writing of my academic papers for quite a while. The feeling you have that the models are grating (quite a choice of a word, btw) is likely because you start to see the deeper patterns in a model’s use of language. I’m not talking about the obvious tells everybody is referring to. I’m talking about the general voice of the model. After a while you arrive at something of an aesthetic exhaustion. It becomes annoying not because it is bad or sloppy, but because there is just so much out there that all has the same voice.
Thanks Michael. Your note reminded me that I wanted to include something in the article on skills and voice-training, but it slipped my mind - just added as an update.
I suspect the LLM style and the monotony of it is part of this. But I think it goes deeper, and touches on how words lead to understanding in our evolved biological brains. At the same time there’s a chance that to many people AI writing is just fine. And this in itself raises some quite deep questions around the process of meaning making through reading …
One of the things I increasingly wonder about is how many of the characteristics and personas of the various models are due to the underlying model itself (ie the initial weights) vs the various post training strategies that are used (ie someone steered the model in one way or another). Or similarly how many of the clear limitations of AI writing are solvable with different post training strategies. Basically are these things fundamental to the AI models themselves, or some layer on top of it that has made the models think this is right. It’s interesting that some things are consistent across model providers and other things aren’t, eg ChatGPT and Claude have very different voices. With time as more models come out, esp open weight ones, I suspect there will be more variation in the post training that allows us to see. I suspect any domain that’s subjective and “taste” heavy will get disproportionate weight from post training, which makes the role of those people doing the tuning especially impactful. And of course as more of the training data becomes AI generated, then the two become harder to distinguish
Uh oh ... did you redact the em-dashes? Or use them as a red herring? It's like the scene from Princess Bride. "He deliberately used em-dashes in Marmoset to fool us into thinking Rabbit was human. But then he knew we were going to assume this so he did the opposite and guided the LLM to liberally use m-dashes. But of course he knew we would assume he was trying to deceive us by thinking he was deceiving us when he wasn't ...."
"Never go in against a Sicilian when death is on the line!"
Uncertainty is a lovely human quality, one of many. Unequivocal certainty, on the other hand, is hubris or ignorance or both. AI, in my limited one-person experience, professes certainty until you point out something it has written that's provably wrong, and then it's politely apologetic. AI isn't embarrassed about its mistakes the way I would be. It hasn't had my awkward adolescence or the striving for excellence instilled in me by my parents, or the trial-and-error learning that has given me opinions about how I want to sound now, which is quite different from how I would have wanted to sound twenty years ago. Uncertainty is but one of the qualities that comes with a beating human heart.
LLMs owe everything they "know" to the collective writings of a vast array of humans (without giving those humans credit or compensation, by the way.) We humans may be intellectual and logical (at our best) but we are also deeply and unavoidably emotional. LLMs, like vampires, suck the blood from the writing of living humans and their writing sometimes appears to be alive, but it's not alive in the way of something human-written. (Don't get me wrong; LLMs are often very useful vampires.) With their encyclopedic access to info and improving grasp of syntax and rules of argumentation, LLMs can be immensely helpful in shaping and refining what we humans write but ultimately, as Claude admitted to me, "I don't have a heart for things to land in."
Our human language is a unique adaptation in the animal world, but it's our animal nature that makes good writing exhilarating. Reading and writing, at their best, are profound emotional experiences; exhilaration is an emotion! Okay, maybe AI can generate a short story that wins an award, even makes you smile like "sunrise over a sink" (wtf?) or maybe human judges are sometimes wrong, and their decisions falter like sunset over a kitchen sink disposal. Mr. Maynard, I'd trust your writing, with or without uncertainty, over LLM output, regardless of how attractive the undead may sometimes be.
Thanks - and well said!
It's a well known bias of LLMs that they prefer their own outputs to those of other models and to those of humans: https://arxiv.org/pdf/2404.13076. so I wouldn't put too much credence in an LLM telling you that a paper written by an LLM is better to a human written paper.
Thanks - hadn’t seen this paper. Of course this is a moving target with the sophistication of agentic frontier AI, but clearly from my experience still an issue. That said there is a rapidly growing community of people who take LLM critique as the gold standard …
Nice post. I’ve used AI assisted writing of my academic papers for quite a while. The feeling you have that the models are grating (quite a choice of a word, btw) is likely because you start to see the deeper patterns in a model’s use of language. I’m not talking about the obvious tells everybody is referring to. I’m talking about the general voice of the model. After a while you arrive at something of an aesthetic exhaustion. It becomes annoying not because it is bad or sloppy, but because there is just so much out there that all has the same voice.
Thanks Michael. Your note reminded me that I wanted to include something in the article on skills and voice-training, but it slipped my mind - just added as an update.
I suspect the LLM style and the monotony of it is part of this. But I think it goes deeper, and touches on how words lead to understanding in our evolved biological brains. At the same time there’s a chance that to many people AI writing is just fine. And this in itself raises some quite deep questions around the process of meaning making through reading …