Jacob Polay with Chloë Farr and Jessica Jack
This post is part of a series on AI and Collaboration.

Transcribing archival documents is an important part of historical research, but a transcribed archive is still just prose. Before a historian can map trade routes and build databases to trace a merchant across ten thousand pages, they must turn this prose into useful data, identifying and tagging the persons, places, organizations, and other relevant entities within the text. Doing this by hand at a useful scale is impossible. Instead, historians most commonly turn to software called Named Entity Recognition (NER), which powers most of the structured historical data you’ve ever queried, from tagged court records to place-name indexes.
NER has evolved through three generations. The oldest systems were rules-based: using sophisticated, frequently hand-curated, gazetteers that matched text against lists and patterns. The second generation transformer models used the context of the sentence both before and after each word to distinguish that word’s meaning. For example, these systems could separate “Turkey” the place from “turkey” the animal by identifying context clues, like reading “Istanbul” or “feather,” in the sentence around the word. The newest generation uses large language models (LLMs), colloquially known as AI, which bring general linguistic knowledge to the task. This generalized knowledge allows these models to perform at a higher standard, providing more accurate results when identifying historical entities. Each generation improves on the last, but all share a limitation in that each model performs only as well as its training data resembles the documents under analysis, and almost nothing is trained using historical documents, especially ones as old as the early modern period.
The consequences of this lack of historical training are significant. To measure NER performance, I ran one hundred pages of OCR’d early modern text through fifteen systems and compared every tag against a hand-tagged gold standard covering four entity types: people, places, commodities, and organizations. spaCy, the free tool commonly used by humanists, missed most of what matters and scored just 20 percent on the standard accuracy measure. General LLMs like Google’s Gemini 3 Pro performed better, reaching 67 percent, but its failures are telling. Gemini hallucinated over 450 false entities, manufacturing patterns that do not exist in the sources. Other typical errors included entities filed under the wrong type entirely, such as reading the “East-India Company” as the place “East India.” These errors are not random noise that washes out at scale; they’re systemic, mis-tagging the same items the same way every time.
Part of the reason those systems created these kinds of errors was that they were not created for historical research. However, looking at these errors from a historian’s lens allowed me to create a better solution for history. Additionally, I noticed that those systems had a tendency to repress those entities which appear only rarely in the corpus. Sometimes, the most valuable entities, the obscure people and the once-mentioned places, are the very heart of where historical discovery is found. Fixing the systemic problems, and surfacing these obscure entities, were my goals in creating my own solution.
Through testing, the evident solution wasn’t to find a better product. It was building one. Over a few weeks working with an LLM coding assistant, I fine-tuned a small open-source model on training data I curated by hand, producing EarlyModernNER. This model, which was trained on a home computer using consumer grade hardware, scored 82 percent, roughly 10 percent higher than Gemini 3 Pro, while filing zero entities under the wrong type, which no other model could do. It runs on most computers and is free to access via Github. More importantly, it is completely open and retrainable, meaning any historian with their own hand-tagged sample of sources can take the model and teach it their period’s documents.
The exact numbers achieved by EarlyModernNER matter less than the decisions behind them. Every design choice was a historical judgment: privileging precision over recall, because a hallucinated entry is worse than a missed mention; selecting entity categories that reflect what my historical investigation required; curating training data as an act of source criticism.
For decades, humanists have been consumers of models built by and for other fields, applying them to our studies with little modification. With the emergence of LLM coding assistants like Anthropic’s Claude, small but mighty open-weight models, and consumer hardware becoming more powerful, the creation of digital tools for historians—once the function of the all too rare funded lab—is now within reach of a single researcher. Off-the-shelf tools may fail to produce useful data with your sources, but that’s no longer the end of the story. It’s the starting point. Historians now have the power to create our own tools, which can be shared with one another to empower the research of the entire field. Every gold standard, every shared dataset, every time a historian publishes open code to GitHub raises the floor for the next researcher, allowing our computational labour to compound on itself the way our research always has with footnotes. Now is the time to build. But even more, now is the time to collaborate.
Jacob Polay is a PhD student in History at the University of Saskatchewan, studying the roles Large Language Models have in the historical method. His current research involves creating an information retrieval pipeline using artificial intelligence tools to unlock the early modern archive at scale.
Chloë Farr is a researcher working at the intersection of artificial intelligence, archives, and digital humanities. Working out of the Open Science Lab at TIB – Leibniz Information Centre for Science and Technology, her research focuses on large-scale text recognition and analysis of historical documents, including newspapers, maps, and archival records. Learn more about Farr’s work on GitHub.
Jessica Jack is a PhD student in History at the University of Saskatchewan, developing applications for Large Language Models in historical research. They are doing so through studying settler land use in late 19th century and early 20th century Saskatchewan.

This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License. Blog posts published before October 28, 2018 are licensed with a Creative Commons Attribution-NonCommercial-ShareAlike 2.5 Canada License.