Jacob Polay with Chloë Farr and Jessica Jack
This post is part of a series on AI and Collaboration.

Transcribing archival documents is an important part of historical research, but a transcribed archive is still just prose. Before a historian can map trade routes and build databases to trace a merchant across ten thousand pages, they must turn this prose into useful data, identifying and tagging the persons, places, organizations, and other relevant entities within the text. Doing this by hand at a useful scale is impossible. Instead, historians most commonly turn to software called Named Entity Recognition (NER), which powers most of the structured historical data you’ve ever queried, from tagged court records to place-name indexes.
NER has evolved through three generations. The oldest systems were rules-based: using sophisticated, frequently hand-curated, gazetteers that matched text against lists and patterns. The second generation transformer models used the context of the sentence both before and after each word to distinguish that word’s meaning. For example, these systems could separate “Turkey” the place from “turkey” the animal by identifying context clues, like reading “Istanbul” or “feather,” in the sentence around the word. The newest generation uses large language models (LLMs), colloquially known as AI, which bring general linguistic knowledge to the task. This generalized knowledge allows these models to perform at a higher standard, providing more accurate results when identifying historical entities. Each generation improves on the last, but all share a limitation in that each model performs only as well as its training data resembles the documents under analysis, and almost nothing is trained using historical documents, especially ones as old as the early modern period.
The consequences of this lack of historical training are significant. To measure NER performance, I ran one hundred pages of OCR’d early modern text through fifteen systems and compared every tag against a hand-tagged gold standard covering four entity types: people, places, commodities, and organizations. spaCy, the free tool commonly used by humanists, missed most of what matters and scored just 20 percent on the standard accuracy measure. General LLMs like Google’s Gemini 3 Pro performed better, reaching 67 percent, but its failures are telling. Gemini hallucinated over 450 false entities, manufacturing patterns that do not exist in the sources. Other typical errors included entities filed under the wrong type entirely, such as reading the “East-India Company” as the place “East India.” These errors are not random noise that washes out at scale; they’re systemic, mis-tagging the same items the same way every time.
Continue reading









