Old OCR text cripples language model training, and FineBooks wants to fix that at scale
Researchers from Hugging Face and EleutherAI have benchmarked 14 open-source optical character recognition models to improve the quality of historical text used for AI training. The project identified dots.mocr as the most effective tool, achieving 97.6 percent character accuracy at a low cost. This development addresses a significant bottleneck in language model development, where low-quality digitization of historical books often introduces errors into training datasets.
Covered by 1 source
- TThe Decoder↗Matthias BastianAug 10