How Can We Minimize the Environmental Impact of Our AI Use?
It's not easy being green when there's little research on the energy consumption of specific AI tasks
by Dan Cohen

* * *
[A version of this piece will run in the Boston Globe as part of the newspaper’s new College Town newsletter on higher education, which I encourage you to subscribe to. My thanks to Brian Bergstein, Senior Editor, Globe Opinion, for reaching out about occasional collaboration between my Humane Ingenuity newsletter and College Town.]
* * *
When I arrived at the Roy Rosenzweig Center for History and New Media at George Mason University in 2001, we hosted roughly a dozen websites on various historical topics. This entire portfolio was running on an old Macintosh that sat quietly in the back of a closet. Several of us young historians persuaded the center to upgrade this modest setup to something more 21st century — a refrigerator-sized rack with maxed-out servers and a row of whirring hard drives.
This may have been too much computer for our purposes. Those websites did load faster — at least that’s what it seemed like — and there was a dazzling array of blinking lights, which were sci-fi cool, but the fans blew a sirocco at anyone who opened the closet door. Lord knows how much power this hot rod consumed.
At the time — and, frankly, in the decades afterward — there was relatively little outrage in academia about the environmental impact of the digital infrastructure we came to rely on. Every university library replaced its card catalog with a web-based platform, every lab ported experimental results from notebooks to databases, every department got a website, every registrar, treasurer, and administrative unit ran on expensive, complex software. In a windowless building at the edge of campus, or, increasingly, somewhere far away from the eyes and ears of students and faculty, were dozens or hundreds of the pulsing, noisy beasts that had turned our history department closet into a sauna. If there were environmental concerns about this sweaty menagerie, they were largely addressed through counterbalancing initiatives, such as solar panels on campus buildings, rather than through the minimization of computational tech.
Only in the last few years, with the rise of artificial intelligence, a technology that unsettled us in many other ways, did the calls to address the environmental impact of computers become loud enough to hear over the din. The growth of data centers dedicated to AI is, of course, far more rapid and rapacious than the proliferation of data centers for websites or web-based software, which makes the pressure to curtail the insatiable energy demands of our modern digital infrastructure understandable. But even before the release of ChatGPT, numerous cavernous data centers arose in Northern Virginia near George Mason; we were just ignorant of their existence, or didn’t care that much.
Now that the issue is more salient, how should academics and anyone else who uses AI think about the environmental costs of this machine labor? If you are committed to a pure anti-AI lifestyle, I suppose you have a clear answer to this question — avoid using AI for anything — and you can sleep well knowing you are not frivolously chatting with bots that are heating the planet. (You could probably be even more pure, however, since you are reading this piece on a digital device that also relies on data centers, those older ones that quietly multiplied before AI; the life of the ascetic is one of eternal self-discipline.) But if you believe, as I do, that there are helpful uses for AI, not for everything but for some things, just as the web had some great advantages, and you also care about Mother Earth, what then?
* * *
Nearly four years after the launch of ChatGPT, we still lack the rough intuitions about minimizing the environmental impact of our AI use that we have for other aspects of our energy consumption. With transportation, for instance, we clearly understand that it’s better for the planet if we walk, ride a bike, or take public transit than if we drive a gas car. By contrast, when I use Claude for some basic tasks, I wonder: Do I really need to be using this cutting-edge model for my pedestrian purposes? Do I need a truck when I could use a bike? If I instead download one of the small free or low-cost open LLMs and run it on my personal computer or my university’s servers, am I actually consuming less energy than Claude uses in one of those dreaded new data centers? Counterintuitively, it could be the case that the most advanced frontier model uses less energy per task than a local model does, in the same way that our dedicated history server from the early 2000s almost certainly used more energy per task than a later version running in the highly optimized cloud infrastructure that now powers much of the internet.
There is precious little guidance for students, faculty, and the public about the lowest-impact AI-assisted approach for specific tasks. Right now much of the discussion about curtailing the cost of AI use is financial, driven by businesses trying to route individual work tasks to the most effective models at the lowest price. This rising cost sensitivity may overlap with environmental sensitivity — after all, lowering payments to AI companies often entails reducing the number of tokens and computational cycles expended on each task — but the chemist, economist, psychologist, and historian have niche disciplinary tasks that might be vastly different from each other, and from those found outside of academia.
So we find ourselves in a situation in which we might want to act locally, but are only able to think globally, focusing on the aggregate impact of AI as seen in the current data center expansion. We have no counsel, other than simple calls for abstention or reduction in use, for our personal or organizational needs.
What we need is something like a Consumer Reports that can point to the most energy-efficient AI model and mode for each academic task. (This could also include suggestions about when AI isn’t needed at all.) The key thing for us to recognize is that tasks vary widely in computational intensity. Take two examples I’ve explored in this newsletter over the last year: natural-language searching of the library catalog and handwriting recognition of historical documents. For the former, the system only needs to understand what a user is looking for, and to translate that into a well-formed search query that is passed to the more traditional digital environment of the library — its database of books, journals, and other resources — from which the main results will be returned. For this task, it seems likely that a very small LLM could be as effective as a large one. We are not asking the LLM to produce its own interpretation or answer, just to act as an interlocutor between the library patron and the collection.
On the other hand, handwriting recognition, especially of older, imperfect documents containing unusual layouts or archaic languages — the kind that one finds in archives and special collections — seems to require considerably more firepower. As I’ve noted before, general-purpose AI models are now able to brute-force their way to understanding these historical documents, but specialized models can do the same or better with less effort. Last year, five computer scientists from Stanford assembled a diverse set of writing that spanned over 2,000 years and dozens of languages, and then tested open and closed AI models for both the accuracy of their transcriptions and the associated (financial) cost.

The researchers then showed that they could produce a very lightweight and effective model of their own, which they called Churro, by fine-tuning an open model (Qwen, one of the Chinese AIs that have made American policymakers and Silicon Valley titans nervous) on this historical dataset of 100,000 pages of handwriting. In the same spirit, four researchers at the University of Basel have been developing a dashboard that color-codes the competency of various AI models on basic humanities tasks, such as extracting bibliographic information from a publication or the correspondents from a letter.

If benchmarks like these added a rough analysis of energy use, we could begin to move toward ecologically sound, computationally efficient, and effective AI use for our teaching and research needs. The energy cost of each model/task pairing might initially be estimated from the size of each model, the number of tokens used for each task, and the length of computation. To supplement this data, universities could demand an ecological scorecard for each model, locally run or hosted in the cloud.
This important work to find the highest-accuracy models at the lowest cost for each of the specialized tasks in each discipline has only begun, but it is an essential step in minimizing the energy usage of AI for academic purposes. It would be a significant improvement over our current, occluded view of our environmental impact.