Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale
Researchers have released Institutional Books, an open-source pipeline designed to clean, deduplicate, and annotate large-scale datasets of digitized library materials. The project includes a collection of nearly one million volumes from the Harvard Library, providing a structured resource for developers and academics working with historical optical character recognition data.
Covered by 1 source
- AarXiv CS.AI↗David Lowry-Duda, Matteo Cargnelutti, Catherine Brobston, Salwa Ismail, Greg Leppert, Amanda Watson, Jonathan Zittrain2d ago