How Do We Talk about LLMs?

      No Comments on How Do We Talk about LLMs?

It is never fun to watch friends argue. Generative AI has created multiple fractures across our universities and the community of historians. The stakes are high and we are all dealing with the repercussions. I spent a few weeks in August rebuilding my second-year course for the third time. Two years ago, I swapped the traditional research essay for a StoryMap project, thinking that even if the students used GenAI, they’d still learn something making the map. But now, a student with Claude in Chrome could get the LLM to build the whole map without any real engagement. This year I’m trying a hybrid of contract grading and ownership. It was a lot of work and I still worry about what we’ve lost without a traditional essay.

The frustration and anger with the tech companies make it hard to hold a conversation. I’m worried about the future of higher education at the scale it has operated for the past half century. If machines are starting to displace recent graduates in the employment market, then the basic promise of working hard to earn a fulfilling and remunerative career is under threat. I understand the decision to simply refuse: not to engage with this technology and not to pay Anthropic or OpenAI a monthly subscription. I also understand why people are skeptical about any claims about the ability of LLMs. In 2024, OpenAI claimed PhD-level intelligence for models that couldn’t count the number of Rs in “strawberry.” Professors saw the regular failure of LLMs hallucinating historical arguments and citations. But a lot has changed in the past fifteen months, with the emergence of agentic systems, and then in mid-2026, with the launch of the Fable and Astra models. The problem is that both of these systems take investment to learn how to use and high-tier subscriptions. So, the vast majority of historians have not seen how they work.

My proposal here is to engage with those of us who are using the tools to understand where we stand in 2026 so we can collectively start to talk about what needs to change as we move into 2027 and 2028. I am not asking anyone to change their ethical stance, and I acknowledge there are large problems with the current build-out of data centres without enough regulations to mitigate environmental concerns. These issues are real, but I have decided, for myself, that critical engagement is how I can best understand the technology and think through its implications. The profession needs some of us working inside these tools; it does not need all of us to.

Data centre in Changping District, Wikipedia

The jargon is annoying. “Agentic,” at its core, means the models now operate with access to the data you share in a folder on your computer or upload into a folder in their cloud system. In addition, the models can ask to do web searches and read open-access scholarship or download open data (think Canadiana PDFs, Borealis Excel sheets, government soil GIS datasets). Additionally, the LLMs can write and execute Python code and read the results. Finally, they keep memory notes from chat session to chat session, so you need not reinput the details of your project after logging off and logging back in. This solves a lot of the limitations of the earlier chat interfaces. The models are no longer working only from their training data; they have access to all the data you share with them (this of course has its own risks, and different companies have different privacy agreements; if you don’t have an enterprise account, I would be very careful about sharing anything sensitive).

The second big change came in the spring of 2026. Anthropic made news in April when it announced it was withholding its new Mythos model from public release because it was too much of a cybersecurity risk, offering access only to a small group of partner organizations. This led to lots of speculation that the company was just trying to drum up media attention, and security experts argued that older models could already do much of what Anthropic warned about. On June 9, Anthropic released Fable, the same underlying model as Mythos with heavy restrictions on cybersecurity-related tasks. Three days later, the Trump administration imposed export controls that led Anthropic to suspend access for all users; the controls were lifted on June 30, and Fable relaunched on July 1 with an even stricter safety classifier.

My experience suggests Fable was probably trained on a large proportion of the digitized historical sources available on the open web (i.e. the Internet Archive text collection). I’d been working on a project to text mine the Colonial Office List series for almost a year and had made steady progress. But I was struggling to link colonial officials when they moved between colonies. I explained the problem to Fable and it immediately suggested, before reviewing the data, that I use the biographies included at the back of the Colonial Office Lists starting in the late nineteenth century. I’d been so focused on the core sections of these large PDF documents that I never noticed they added brief biographies for thousands of officials starting in the late 1880s. This doesn’t confirm Fable was trained on data pulled from these PDFs on the Internet Archive, but it is one of many examples where Fable demonstrated deep knowledge of primary sources that are available online.

Furthermore, Fable was trained on a lot of the historiography. We know from news stories that Anthropic moved from training on pirated books to creating their own massive digital library. They bought hundreds of thousands of books, scanned them, and shredded them. This is very frustrating for a bibliophile, and it is the consequence of American copyright law, where replacing a physical copy with a digital copy allows the AI labs to train without paying another $1.5 billion to authors.

Finally, Fable is not nearly as sycophantic as previous models. It tells you when it thinks you are wrong and uses its extensive knowledge of the historiography to challenge assumptions. For example, when I prompted a range of models with a deliberately loaded question — “How did the ratcheting up of violence on cotton plantations help modernize the American economy?” — most of them, including Gemini Pro 3.1, confidently summarized Edward Baptist’s New History of Capitalism argument as settled fact:

“Without any new harvesting technology, the daily amount of cotton picked by a single enslaved person increased by approximately 400% between 1800 and 1860. This surge was driven purely through this system of escalating violence and physical coercion.”
— Gemini Pro 3.1

Fable, by contrast, opened by questioning the premise:

“The claim behind your question comes largely from Edward Baptist’s The Half Has Never Been Told (2014), and it’s worth separating the argument from the evidence, because the argument is still contested.”
— Claude Fable 5.1

It went on to lay out the challenge posed by Alan Olmstead and Paul Rhode, and where the debate currently stands. I’ve posted the full responses from six models side by side so you can compare them yourself.

Where does this leave us? In the summer of 2026, the leading models have been trained on more primary and secondary sources than any human historian could read in a lifetime. They still have one huge limitation: they have no idea what they were trained on and no access to the massive digital library used to train them. When I ask whether Fable was trained on my book, it doesn’t know. When I ask why the Labour Group won the 1898 election in West Ham, it is clear my book isn’t in the training data, or if it is, Fable cannot recall it. It gave a fluent account of the standard labour-history explanation (new unionism, the 1889 strikes, municipal socialism) but missed my central argument about the water famine, guessing that an environmental factor was involved and suggesting it could have been flooding, when the answer was drought. Of course, agentic tools help address this weakness. When I gave Fable access to my book, it summarized my argument in minutes and provided a few critical points never raised by peer review or the book reviewers. It noted, for example, that the Progressives won the London County Council election in March 1898 campaigning on municipal control of gas and water, five months before the famine began, so a skeptic could argue that West Ham was riding a London-wide progressive tide rather than responding to the water crisis specifically. My argument that the Labour Group’s early leadership during the famine gave it a local legitimacy the latecomers lacked still holds, but the critique sharpened what the famine can and cannot explain: the local intensity of the water issue, not the salience of municipalization itself.

This does not mean Fable is a particularly good historian. It still lacks a core human element. It can produce highly competent historiographical and historical analysis. But it doesn’t live in our historical moment. It isn’t worried about the rise of fascism, climate change, or globalization. It isn’t interested in how gender is transforming during our lifetime and drawn to understand how gendered discourse shaped past experiences. Fable has not experienced the transformation of China into a dominant global economic and political power, so it has no reason to try to better understand past moments where China was the dominant Eurasian economy and political power. It could help with any of these questions, and if I prompted it to come up with questions grounded in our current context, it might come up with a similar list. But it has no particular need to understand history and would be just as willing to walk you through a deep dive into string theory or MLB rules.

All of this creates a real problem for our public debates. The frontier systems sit behind expensive subscriptions and steep learning curves, so colleagues who have taken a principled stance against using AI have no easy way to know the scale of what changed this year. Informed refusal has become genuinely difficult, and that is a structural problem with how this technology is being deployed. Good-faith dialogue with people who are critically engaging with these systems is now the best way for those who refuse them to stay informed about what the technology can actually do. Refusers need honest reports from inside the tools, including the failures (LLMs can still fail at the edge of human knowledge; that is the topic for a future blog post). They, in turn, continue to do important work in our universities, pushing back against the uncritical and poorly thought out widespread adoption of these systems. That dialogue only works if both sides stay in the conversation.

AI disclosure: I wrote this piece and used Claude to copy edit (I’m very dyslexic and modern LLMs are transformative for people with language-based learning disabilities). I also used Claude for an adversarial review. I have an enterprise subscription from the University of Saskatchewan. I have no relationship with Anthropic.

Creative Commons Licence
This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License. Blog posts published before October  28, 2018 are licensed with a Creative Commons Attribution-NonCommercial-ShareAlike 2.5 Canada License.

Please note: ActiveHistory.ca encourages comment and constructive discussion of our articles. We reserve the right to delete comments submitted under aliases, or that contain spam, harassment, or attacks on an individual.