The AI-Enabled Boom in Document Transcription

      No Comments on The AI-Enabled Boom in Document Transcription

Chloë Farr with Jessica Jack and Jacob Polay

This post is part of a series on AI and Collaboration.

Colour photograph of old hardback books on a shelf.
“Old Books” by jarmoluk via Pixabay. Free for use under the Pixabay Content License.

For many people who use archives, the accessibility of archival material is a notable difficulty. Archivists and librarians have long been turning to software to help with that accessibility. One of their main tools is Optical Character Recognition (OCR). This type of software turns images of text into digital characters, which can then be accessed by anyone with an internet connection and can be easily processed and analyzed outside of an archive. But OCR started as a corporate solution to corporate problems, which meant it was not very good at dealing with the differences present in archival materials. Archivists have been trying to address this with bespoke technology, but the last few years have presented a new way to deal with the problem.

OCR enabled by AI vision language models (VLM-OCR) has recently exploded as a research topic in multiple fields including computer science, digital humanities, library information science, linguistics, and beyond. This new movement in OCR began because AI developers hit a wall. They had trained their models on all available data on the web so they tried using AI-generated data instead, but the results were problematic, inconsistent, and risky. To solve this, they turned to physical documents to broaden their training base. However, they needed better OCR to make this possible. The paper by AllenAI that opened the field, “olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models,” says as much in its opening lines: PDFs hold enormous volumes of “novel, high-quality” training data, but their diversity of formats and layouts makes that content hard to extract faithfully.

AllenAI’s statement highlights some of the motivations behind this technology. Many AI companies are now developing OCR models and releasing some of them openly in order to encourage adoption, build user communities, and establish their tools as part of the wider document-processing ecosystem. Open source models—whose source code is publicly available, freely licensed to use, modify, and redistribute—can reduce the cost of access to high-quality OCR. They enable galleries, libraries, archives, and museums (GLAM) institutions and researchers to run, evaluate, and adapt tools on their own infrastructure. However, this does not eliminate commercial competition: companies can still charge for hosted services, enterprise support, specialised fine-tuning, and licences for larger-scale commercial use. For example, Datalab’s Surya makes its tool available under a semi-open license. The ability to fine-tune the tool’s training is free for research, personal use, and smaller startups but they require broader commercial licensing for other users. While useful in some respects, Surya is not made for GLAM and humanities research and thus represents a limited use case of these semi-open corporate solutions.

This kind of for-profit model system can also be seen in Transkribus, one of the most widely used OCR services for humanities research. This software is not fully open source, instead running on a credits-based system where the first 50 credits are free and then the cost increases. Customers can fine-tune or “train” their own models for improved transcription for their collections. Typically fine-tuning is most helpful on homogenous collections with a high volume of documents. However, the company retains the underlying model for themselves, including the clients’ documents used for that development. This restricts the portability of the model outside of the Transkribus environment. Alternatively, Adobe PDF reader is a widely accessible OCR engine requiring little technical knowledge, but is available with a subscription fee, and performs poorly on historical documents. While these are but two examples, the expansive use of these models demonstrate the demand for OCR, but also the simultaneous need for this OCR to become more accessible and sustainable. It is in GLAM’s interest to find ways to use OCR models without relying on for-profit services as open source models ensure ownership of data while minimizing overhead expenses for consistently underfunded institutions. This independence can be achieved through capacity- and knowledge-sharing, and collaborating on OCR processing to remove redundant work.

The outputs of many OCR technologies were developed for training LLMs, but that are frequently of little use in the humanities. They can be used by people with specific data mining and data science purposes, like researchers doing Named Entity Recognition. But these are still rather niche, and the majority of archival researchers and historians who want to access these OCR outputs would instead benefit from the document processing of OCR resulting in searchable archives. This is another area where AI-driven OCR is helpful, as it has the capacity to easily transform OCR outputs into specific and bespoke formats for the needs of the researchers who are using these outputs.

In my work as a GLAM researcher situated in libraries, I focus on making VLM-OCR accessible and useful for research and archival use. I’m frequently asked “What’s the best VLM-OCR model right now?” Before March 2026, it was usually pretty clear. The models were quite uniform, handling the same type of documents, providing the same output formats but with different levels of transcription accuracy based on the source document’s language, scan quality, and text layout. Now, each model has their own distinct strengths. It is a welcome development that models are no longer competing on accuracy alone. For example, Hunyuan OCR provides coordinates for each word on the page, which in turn enables people to search for the word and see it highlighted right on the page. ChandraOCR-2 also does a great job of transcription, and it can detect images inside a document (photos, art, graphics, etc.) and keep them separate from the surrounding text, along with writing short descriptions of what’s in each one. Surya OCR is a very small model, meaning it can run on lower-quality hardware, and transcribes at a higher speed while occasionally sacrificing accuracy. Differentiating by strengths  eases the burden on users, who no longer have to chase a 0.1% accuracy edge and can instead pick whichever model suits their purposes. Regular users seem to develop an intuition for this. It comes with experience, gained by testing different models across a range of document types and matching them to what the user needs from a transcription. In this sense, collaboration between these experienced researchers in the space is key to helping everyone access the models that best suit their needs.

What emerges from this shift is an ecosystem of increasingly complementary AI-enabled OCR models, whose real value depends on the researchers who know how to use them. With these models, institutions do not need to bet everything on one company’s roadmap or pricing model. They can now run the software on local hardware that ensures data ownership stays with the institution. And because the tools are open, GLAM professionals can pool their expertise built through hands-on testing, matching models to materials, and sharing their work across institutions. The models themselves are now good enough that further gains will be small and specialized. What needs improving is how we use them together. The people making these documents accessible need a shared and growing toolkit, built with and for each other to use, made easier by the support of LLMs. For chronically underfunded archives and libraries, that collaboration is worth as much as any accuracy gain. Better OCR is worth having. Building the capacity of archives and libraries to serve the people who rely on them is worth more.

Chloë Farr is a researcher working at the intersection of artificial intelligence, archives, and digital humanities. Working out of the Open Science Lab at TIB – Leibniz Information Centre for Science and Technology, her research focuses on large-scale text recognition and analysis of historical documents, including newspapers, maps, and archival records. Learn more about Farr’s work on GitHub.

Jessica Jack is a PhD student in History at the University of Saskatchewan, developing applications for Large Language Models in historical research. They are doing so through studying settler land use in late 19th century and early 20th century Saskatchewan.

Jacob Polay is a PhD student in History at the University of Saskatchewan, studying the roles Large Language Models have in the historical method. His current research involves creating an information retrieval pipeline using artificial intelligence tools to unlock the early modern archive at scale.

Creative Commons Licence
This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License. Blog posts published before October  28, 2018 are licensed with a Creative Commons Attribution-NonCommercial-ShareAlike 2.5 Canada License.

Please note: ActiveHistory.ca encourages comment and constructive discussion of our articles. We reserve the right to delete comments submitted under aliases, or that contain spam, harassment, or attacks on an individual.