← Back to Model Beat
Open Source·2d ago·all news from August 20, 2026

Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale

Researchers have released Institutional Books, an open-source pipeline designed to clean, deduplicate, and annotate large-scale datasets of digitized library materials. The project includes a collection of nearly one million volumes from the Harvard Library, providing a structured resource for developers and academics working with historical optical character recognition data.

Covered by 1 source

  • AarXiv CS.AIDavid Lowry-Duda, Matteo Cargnelutti, Catherine Brobston, Salwa Ismail, Greg Leppert, Amanda Watson, Jonathan Zittrain2d ago

Related stories

Open SourceMOSS-VL Technical ReportAug 18Open SourceEngineering Signals of Human-AI Collaboration in the Agentic Coding Era: A Longitudinal Analysis of 33,228 Pull Requests from vLLM and SGLang with Implications for Biomedical AI Agents and Bioinformatics Pipeline DevelopmenAug 17Open SourceWyvern: An Agentic Framework for Generating Grounded Multimodal ReportsAug 17Open SourceThe Working Set of a Coding Agent: Coherence Debt in Repository-Scale TasksAug 18