Jacob Polay with Chloë Farr and Jessica Jack
This post is part of a series on AI and Collaboration.

Researchers often face a dilemma when working with digital archives. Once they establish their research questions, they scour the web for digitized archival material, often finding thousands of sources. Next they are faced with the daunting task of turning this diverse array of data, composed of tables, ledgers, letters, diaries, bills, and government acts, into one set of relationships that can answer historical questions. Algorithmic transcription [OCR HYPERLINK] and entity tagging [NER HYPERLINK] help solve the first two steps of creating these relationships. These methods allow the computer to read the sources and organize their information into relevant entity groups to answer the researcher’s questions.
However, these helpful technologies still leave one crucial hurdle that must be overcome. A corpus of 20,000 documents is still 20,000 disconnected files. A search bar can only find specific words or phrases. Context is missing. It cannot answer the questions historians actually ask about the relationships between actor and entity: who traded with whom, who went where, what replaced what, and what came from where. Answering these questions at this large archival scale is much easier and faster with machines, but requires teaching the archive to hold its knowledge the way historians do. Three connected technologies now make that possible, made even more accessible with the use of an LLM.
The first technology is one that humanists have been building toward for decades. Knowledge graphs store information not as tables but as relationships: entities become nodes, and the connections between them become typed edges: sugar GROWN_IN Jamaica, sugar HARVESTED_BY enslaved Africans, sugar EXPORTED_TO London. This technology may sound familiar. It’s the same structure powering software like Uber, AirBnB, and eBay, and is also the foundation of Linked Open Data which is the basis for Wikidata, Geonames, the LUX: Yale Collections, and others. These open sources publish entity-relationship data with shared identifiers so that the historical spelling “Barbadoes” in one online collection resolves to the same island as “Barbados” in others. A knowledge graph is essentially this idea at the scale of a single project: a map of every documented relationship in your corpus, queryable in milliseconds.
Continue reading








