14 Comments
User's avatar
Richard Domurat's avatar

Interesting read. This has been very similar to my experience as well, so I’ve generally had a similar conclusion about the ability to augment but not replace researchers. But what’s still hard for me to see is, with the rate of change, does that categorically change the conclusion, or does it just become a better assistant. Eg if we fit a trend line to model capabilities, at what point does this conclusion no longer hold, if at all?

Andrew Maynard's avatar

That is THE question! It's also one - while I don't have a good answer - that drives the question of what the purpose of research, discover, curiosity, enquiry, learning etc are in society — and to us as individuals. If it's just ultimately the creation of wealth and material health and wellbeing at speed and scale, I think you can extrapolate to a near future where AI increasingly takes the lead. But when you factor in other dimensions of meaning and thriving as a human, I think things get more complex,

Joel Hughes's avatar

“Maybe they can get it to produce stuff that they think is good.” So true. This is my concern as I try the same idea. I created a project and skill trained on select examples of my writing to attempt to get Claude to follow my writing style. It helped. I made a “tabula rasa” project with instructions to start from scratch to attempt to avoid context contamination (or work in one domain “leaks into” another somehow: I didn’t want ideas for how to teach the paper content to my undergraduates for example). I wrote instructions to identify and challenge assumptions inherent in my prompts, etc. the result was impressive “critical thinking” that led to revisions and improvements. I never expect Claude to write “like me” and anticipate that I will always need to remain a developmental editor and “insert my personality.” Also, I don’t trust its research in my area because it can’t past the paywall and there will always be some selection effect on the search space (never truly systematic: but most people wouldn’t notice). After 2 days (July 2-3) the result is far from finished, and I have a nagging fear that the thesis and argument are shallow and misguided because the topic is “adjacent” to my expertise (I’m a clinical health psychologist with a longstanding interest in the implications of cognitive science for pedagogy - yet I am not a true cognitive scientist). I will need peer pre-review for the “sniff test,” or whether the paper makes a novel non-trivial contribution or is a pile of carefully arranged dung.

Andrew Maynard's avatar

Very similar to my experiences. Fable is definitely a step up, and is surprisingly good at sniffing out primary sources. But I've learned not to trust its own internal peer review agents as much as I would a human as they read like an LLM and not a human — meaning they can think something is wonderful that would be desk-rejected out of hand.

Joel Hughes's avatar

Unfortunately, in academics, there are so many self-appointed emperors and so few clothes. I worry that “Maybe they can get it to produce stuff that they think is good" will result in people submitting stuff that is not very good, and that reviewers and editors will not notice. I mean no disrespect to disciplined, serious researchers, but we all know that some people seeking a 1-shot "research paper" are playing the research game for incentives or out of a lack of self-awareness. For example, I am tired of well-meaning coauthors "writing the discussion" of an empirical paper I circulate for input, only to find their AI-generated slop is just bad. I wonder if thoughtful vs. poor use of AI will start to reveal the Dunning-Kruger effect.

RL Lindsey's avatar

Is there a means to estimate the carbon footprint of this paper? (Unexpected irony, there.)

Andrew Maynard's avatar

Tricky as the estimates of energy use per token out there are unreliable and highly model-dependent. Suspect not insignificant though — although what is even harder is equating energy use/carbon footprint to value of use ...

Eric Jensen's avatar

Andrew, this is really interesting- thanks for sharing your experience. I have sometimes thought of these models as analogous in some ways to a student: smart, eager, a little inexperienced in context and conventions of the field. If you imagined this as a dissertation you’re supervising, where is it at this point? Ready to go to the committee? Or are you sending the student back to keep working? Maybe the analogy doesn’t work here, but your laborious engagement with a written work produced by another entity (not quite sure what word to use there) reminds me of that process.

Andrew Maynard's avatar

getting closer and I use this analogy a lot, but there are differences. Frontier LLMs can research deeper and wider than most students, and make connections between seemingly disparate areas frighteningly well. But they lack formation — the embodied and experience-based ability to make sense of and add value to work in a human/social context — and while they can emulate what they think is effective communication they are missing a secret sauce that, I suspect, comes in part from what it feels like to read and engage with material. With this exercise it felt like I kept sliding around the goal with Fable — approached it fast, but then it kept sliding around where I wanted it to be. I see this will some students, but the ones that succeed pass through the slippery barrier!

VicW's avatar

What was the token $ cost incurred?

Eric Jensen's avatar

Anthropic’s docs say “Claude Fable 5 is priced at $10 per million input tokens and $50 per million output tokens, with the existing 90% input token discount for prompt caching.” So if we take the input and output numbers in the table at face value (assuming all of the input is unique) this would be $10 x 33 + $50 x 3.7 = $515. And perhaps that 359M of “cached content re-read” has to be accounted for at the discounted rate, which would be another $359. So ballpark $875?

Andrew Maynard's avatar

Ouch! I was running this on the Max plan with Anthropic and at a very rough estimate cost around $100 — so far below this, but also very hard to parse out. Would be interesting to run the same paying directly for token use.

VicW's avatar

The big unknown with AI today is what kinds of applications will be cost effective. Until the AI industry has reached some kind of self-sustaining economic status, and one can predict the future cost of using it, one shouldn’t get too dependent on its availability.