Humane Ingenuity logo

Humane Ingenuity

Archives
About
Dan Cohen
Subscribe
September 2, 2026

Scholarship Will Soon Be Much Harder to Save

It's not easy to be confident in a body of knowledge that has phantom limbs

by Dan Cohen

Gold-rimmed glasses lie on a handwritten letter.
Spectacles, Digital Collections Repository, University of North Carolina at Chapel Hill, via DPLA

* * *

[This is the final piece in a miniseries on finding the right line between human thought and AI assistance, focusing on the stages of scholarly work from initial ideas through the research process to publication, although I believe much of this discussion is applicable to intellectual work beyond the academy. The miniseries includes an introduction, an essay on the origin of new ideas, a piece on analyzing evidence and data in the early stages of research, and a careful look at the writing process. In this last issue, I examine what we need to save, post-publication, to ensure that we are confident in the body of scholarship, and can cite and reuse critical elements of it.]

* * *

Two of the most important responsibilities of research libraries are twinned: preservation and access. We save old and new sources of scholarship, and devote considerable effort to ensure they can and will be used decades or centuries in the future, for the benefit of students, faculty, and society. Until a half-century ago, most of these materials were physical, and could be seen on library shelves; since then an increasing proportion of them are digital, and exist in an invisible, virtual space. But we have a decent sense of these materials too, as they are mostly digital surrogates of the older sources: for instance, electronic journals and articles instead of paper ones, ebooks instead of print books. Critically, all of these objects, whether they are made of atoms or bits, are connected through the web of scholarship — the references, quotations, criticism, metadata, bibliographies, and other links that create an intertwined and robust body of knowledge. If you want to understand an academic work in greater depth, you can follow the trail of its connections, and access the various components that went into its construction.

Over the last few decades, this body of knowledge has started to add limbs. Secondary materials associated with the core scholarly outputs, especially data, have become important targets for preservation and access. In most grant-funded quantitative research in the natural and social sciences, it is now a requirement that the grantee retain and make available the data produced by observation, experimentation, or collection. This record allows other scholars to check the validity of any conclusions, and enables subsequent uses of the data, for instance in systematic reviews and meta-analyses. (It is also a hedge against academic fraud, although the data can of course be fabricated or statistically skewed.)

Data is much bigger and harder to preserve and make accessible than a common digital object like a PDF, and this has made the work of the research library increasingly challenging. Large data sets are often not even deposited in the library, where they would be adjacent to relevant writing and analysis, but elsewhere, such as on the servers run by the university’s research computing unit, on a shared data system for a field (like ICPSR), on a journal’s website, or, heaven forbid, on a personal or lab drive. This creates a rockier landscape for preservation and access, and makes it harder to guarantee that a scholarly work and its raw materials can be found together. “Preservation and access” sounds simple, but in each generation, we must ask what to preserve and how to make it accessible in perpetuity, which is not so simple.

* * *

Now comes the emerging AI-assisted workbench for researchers, which has not a single additional limb but more tentacles than an octopus, some of them hard to grasp. Take Claude Science, currently in beta. For those who haven’t used it, Claude Science is a specialized app that combines Anthropic’s LLMs with “connectors” (a plainspoken rebranding of the Model Context Protocol, or MCP, covered in this newsletter a year ago) that can search chemistry, biology, and physics journals, as well as other scientific databases; dozens of “skills,” detailed recipes and instructions that tell Claude how to operate on scientific data and other materials (like MCP, skills were developed by Anthropic, but are now an open standard); and the ability to add extra computing power from a campus server or the cloud. In other words, Claude Science is a custom AI-assisted rig, or disciplinary workbench, a merger of formerly distinct elements in scholarly practice — resources, data, articles, reference works, methods, computation — into a single package, tailored to a specific academic field. (Disciplinary workbenches aren’t confined to the sciences; here’s a group of historians building one for history.)

With a discipline-focused design, rather than the generalized, open-ended interface of a standard-issue AI chatbot, workbenches could very well emerge as the AI assistants I have been looking for in this newsletter miniseries: rigorous tools that can help us with highly specialized scholarly explorations, while centering human intelligence and respecting the existing scholarly record — indeed, constantly working with that record through connectors to library sources.

But like a cephalopod, the assemblage of the disciplinary workbench is also elusive and a bit scary. For instance, under the hood, Claude Science allows a researcher to add AI agents/personas of any kind as virtual assistants in the workspace — e.g., “give me a postdoc who is excellent at statistics, but is reserved unless I’m making a terrible assumption in my stats, at which point they will become surly and correct me with abusive language.” One built-in agent, evidently conjured from the mind of Philip K. Dick, is called The Reviewer, which can check for errors as the work proceeds, making sure, for instance, that no citation has been hallucinated. There’s not an agent (yet?) called The Writer, to compose your whole paper, but as I’ve discussed recently, it’s a slippery slope. And even the idea of The Reviewer and other AI assistants and interlocutors may, of course, incline researchers to reduce the time they spend with human colleagues, leading to the “scientist as conceptual artist” problem I identified back in 2024.

The fusion of agentic AI, compute, research protocols, articles, and data is a much more complex, heterogeneous scholarly object than a book, journal, or PDF. It’s also much less fixed than those traditional objects. One curious aspect of LLMs, which sit at the command center of the disciplinary workbench, is that they are nondeterministic. When you engage with existing scholarly applications, such as an index-based search tool for journals or an analytical application such as Mathematica, these tools return the same response to the same input every time. This is not true for LLMs, which will give you a slightly different answer each time — a rather odd characteristic we are still getting used to.

This fluidity is acceptable when planning a vacation or composing an email message, but it is highly problematic for the scholarly record. Accurately preserving scholarship that used AI assistance requires far more work. For example, we might want to note which AI model was used by the researcher and store a transcript of their interaction, but is that really enough information? Do we also need to preserve the model itself? If so, how do we do that, given the enormous scale of even the smallest current models, not to mention the massive models to come? Moreover, since the outputs of LLMs are nondeterministic — in my experience, Claude’s search of academic literature returns a different slate of articles for identical research queries — simply recording the model used does little to preserve how it acted upon data, text, or other materials. Worse, disciplinary workbenches might have multiple agents running parallel processes for minutes or hours, rather than a single chat stream. How do we begin to understand and preserve what happens during these processes? If we are struggling to reconstruct what a swarm of AI agents did during a hacking event this summer, will the record of a university research project be at all legible in the coming years?

A list of the connectors available to the LLM is similarly amorphous: it gives us only an overview of the vast resources a scientist may have consulted, and not a concrete list of all of the materials scanned, summarized, or synthesized. Hopefully individual journal articles will be referenced in a more traditional fashion, i.e., in footnotes and a bibliography, but that is not guaranteed. More promisingly, we could save the skills documents in a preservation package, since they are relatively slim text files, and we have some experience saving similar files, such as with Protocols.io, a service that stores the rigorous processes a researcher might use in an experiment.

One thing is clear: the preservation of the scholarly record is going to require more intrusion into the researcher’s work and digital environment. If we continue to save only the end products of scholarship, such as an article or book, and associated artifacts like data sets, will there be an erosion of trust in the scholarly record given how many other digital tools played a role in the research? How can we discover and diagnose researcher errors when we can’t access the distinct blend of skills, connectors, data, resources, and AI models behind them? If the disciplinary workbench becomes a black box we can’t see into or understand, we will have trouble reconstructing the scholarly process and its conclusions.

* * *

Interestingly, the makers of Claude Science seem to share this worry. In addition to The Reviewer, the software pays attention to the “provenance” of each research element or output. According to Anthropic’s documentation, Claude Science “saves results as versioned artifacts with a full provenance record” so others can inspect the code used in the artifact’s creation, the environment it ran in, a plain-language description of what was done, and the LLM conversation history. Much more could be done to enhance the scholarly record generated by disciplinary workbenches. Today’s AI models can summarize and store longitudinal records of a researcher’s interactions, and those could be accumulated into an archival manifest for an article or book, then ingested by a university’s institutional repository. Those records, like skills, could be human-readable text files, which would make future access easier. Anthropic should give The Reviewer a buddy: an agent I would call The Librarian.

But long-term access to the other AI elements of scholarship is likely to be considerably more complicated. Will it even be possible to recreate a scholar’s virtual workbench in 50 years? It might not be possible in five. Before the advent of LLMs and related AI tools, we had two approaches to preserving and providing access to software: emulation and migration. You could either create a modern system that runs old applications — as seen in the Internet Archive and the Software Preservation Network — or regularly migrate formats forward in time so they keep working, at the potential loss of some functionality and fidelity to the original. But will we really be able to emulate or migrate gigantic AI models that constantly morph over time, especially if they are not open and available to download?

And if we can’t preserve these models and other AI tools, will there be any way for a future researcher to do the kind of scholarly forensics that are still possible today? Or will our ability to reconstruct important breakthroughs falter because some of the limbs of a once-solid body of knowledge have turned into phantoms?

* * *

Read more:

  • February 3, 2026

    Where Should Scholars Draw the Line on AI?

    Between the poles of Zero AI and AI For Everything lies a vast, poorly mapped middle ground

    Read article →
  • July 14, 2026

    Perils of the New Armchair Scholarship

    Like nineteenth-century scholars who wrote about the world without deeply engaging with it, academics who lean too heavily on AI risk becoming paper-thin

    Read article →
Don't miss what's next. Subscribe to Humane Ingenuity:
Older → A Major New Grant for a Library-Led Approach to AI and Books