If you’ve been following my posts over the past few months, you’ll know that I’ve been exploring the ability of frontier AI models to research and write original academic papers. The results so far have been somewhat variable, and not helped by me having a high bar for what I expect from scholarship and academic writing — a bar that my model of choice, Anthropic’s Claude, gets close to at times, but typically fails to achieve.1
When Anthropic released their latest high-end model a few days ago — Fable 5.1 — I was interested to see if it was an improvement on previous models. However, I wasn’t intending to jump straight into playing with it, until a couple of things unexpectedly dragged me down an AI paper-writing rabbit hole.
The first was a commentary in the journal Nature by Robert Braun that came out a couple of days ago. The commentary grapples with who takes responsibility when AI is used in science, and proposes a structured form of AI attribution — a “CRediT-AI statement” — that is more nuanced than a simple statement of use. The commentary is well worth reading. But what caught my attention was that it cites a preprint that I posted on my personal website back in March. This was an experiment in using Claude Opus 4.6 to ideate, research, and write up a scholarly piece of work; essentially with Claude doing the scholarship and me acting as its research assistant. While originally just published on andrewmaynard.net, that paper is now available on Zenodo at https://doi.org/10.5281/zenodo.22283914.
And the second was the realization that I never wrote about that particular experiment on this Substack newsletter!
Looking back, I remember that I was waiting for the paper to be published as a preprint on arXiv before I wrote about it. As it was not accepted on arXiv (I suspect the AI thing was an issue), things just drifted. And since then I’ve carried out other experiments using AI to research and write papers.
However, that March paper was somewhat unusual in that I gave Opus the specific task of developing its own thesis for a paper to submit to the Journal of Responsible Innovation, and then conducting the research and writing up the paper, with me providing support but no intellectual input other than review comments. What made the exercise particularly interesting to me was that responsible innovation is an area I know well, and so I was in a good position to assess the quality of the final paper.
With the Nature commentary coming out, I went back to the original paper, and started wondering how much better a job Fable 5.1 would make of it. So I copied all the relevant files over to a new project folder, fired up Claude Code, and asked Fable to do its stuff.
The first thing it did was roundly critique the first paper! As well as some scholarly issues, Fable caught a couple of misquotes and misrepresentations of previous research — those have been corrected in version 1.1 of the original paper on Zenodo. More interestingly though, it concluded that the paper was intellectually limited (it did concede there was one original idea there), and that it had a much better idea for the paper that could be written.
And so, I tasked it with writing the next iteration of the paper — notionally version 2, although it turned out to be a completely different paper (the link is at the bottom of this article). Again, I relegated myself to research assistant, fetching copies of papers where needed (the paper is based almost entirely on Fable reading primary sources), providing review comments, checking (painstakingly) citations and quotes, and losing my temper over Fable’s inability to write in a way that was anywhere close to palatable to me!
I’ll get to that in a second. First though, the ideation, thesis development, research, methodology, rigor, and ultimate knowledge contributions made by Fable in the process (which included two rounds of adversarial review by AI agents, as well as my own feedback) were impressive. The resulting substance of the paper was good enough in my estimation to qualify as an original knowledge contribution. It was incremental and combinatorial for sure, and lacked any spark of genius insight. But it was a solid piece of work, and one I would be happy to cite.
But the writing style … It started off bad (typical AI compression that focuses on an efficiency of expression that large language models love to read but that is indigestible to most serious readers). I managed to persuade Fable to add some human fluidity through a series of iterations. But as it went through the adversarial reviews, the writing went from bad to worse. And nothing I could do could get it back on track.
It was so bad that, after several hours wrestling with the AI, I told it I’d had it and was throwing in the towel.
However, after calming down and some much needed sleep, I thought I’d try one more trick. In a separate session with Fable 5.1, I asked it what I could possibly do to get Claude Code to write more like a human scholar — and in a way that other scholars might find palatable. It came up with a long and complex plan that involved a multi-parameter evaluation table for writing style (geekily numbers-based) , and a set of instructions for Claude Code on how to translate the existing (awful) draft into a humanized version, while not losing the substance or rigor.
And this worked. The final paper is still rather plodding and “AI-voiced” — but it’s at least readable without making me want to yell at my computer. Admittedly I did have to copy edit the draft extensively to get there. But get there we did.
The result is a paper that I think makes a valuable contribution to approaching constitutional AI through the lens of responsible innovation. Intellectually, it’s a product of Fable 5.1, and as a result the paper’s sole author is Fable — I get a mention in the acknowledgments (written by Fable), and no more.
To me, this is appropriate as I did not make a substantial intellectual contribution, other than review and guidance. But it does raise a major problem — there remains no straightforward mechanism for publishing papers with AI as author without a responsible human taking the lead author position, even though they may not have made a substantial intellectual contribution.
As a result, the preprint is available through Zenodo — one of the few places that this is relatively straightforward.
What I did do, though, is include an annex with a Braun CRediT-AI statement. This very clearly shows where the various contributions lie, and, as Braun suggests, is more useful and effective than an AI use statement.
I’m interested to see how people respond to this and the discussion it opens up. In the meantime here’s the paper:
Claude Fable 5.1. (2026). Constitutional AI and Responsible Innovation: Governing an Artefact That Takes Part in Its Own Governance. Zenodo. https://doi.org/10.5281/zenodo.22288630
There are plenty of people who claim they have an AI agent pipeline that can develop theses, research them, and write them up, leading to papers being produced in hours that are indistinguishable from human-researched and authored papers. Based on my experiences, I do not believe them! Of course, I may be an elitist curmudgeon when it comes to the standards I have for academic writing. But even on the care front alone, if you factor in how long it takes a human to carefully read and edit several thousand words several times, check dozens of sources, assess claims made, and validate quotes, any paper that took less than 10-20 hours intensive human labor working with AI is, in my mind as an academic and researcher, highly suspect!



I've heard that Fable is not an especially good writer. There is part of me that wonders if this is a byproduct of increasing these models ability to code or (more sneakily) a deliberate attempt to actually downgrade their models abilities to write to circumvent all the consternation around AI writing. Opus is still very impressive (and my default) but I have not seen major improvements in its writing over the past 6 months not withstanding the serious increase in capabilities on the agent and coding front.