Models & Research

Old OCR text cripples language model training, and FineBooks wants to fix that at scale

· August 10, 2026
Old OCR text cripples language model training, and FineBooks wants to fix that at scale

What changed

The FineBooks project from Hugging Face and EleutherAI tested 14 open-source OCR models on a dataset of over 2,000 pages from historical books. They found that the best model, dots.mocr, achieves about 97.6 percent character accuracy while processing at a cost under two dollars per thousand pages. This level of accuracy is sufficient for generating AI training data but not yet reliable for scholarly transcription work.

Why builders should care

Old OCR text is a hidden factor that degrades AI model training quality. Historical books often rely on outdated and error-prone OCR, which introduces noise into training data and weakens language models. Improving OCR at scale helps eliminate this bottleneck and can lead to cleaner and more accurate AI training datasets. For anyone building AI systems that depend on historical or public domain text, FineBooks points to a path for better input data with manageable costs.

The practical takeaway

If a project or startup relies on digitizing historical text collections to build datasets, the traditional OCR accuracy limitations are a real operational headache. FineBooks shows that cutting-edge open-source OCR models now make it economically feasible to upgrade these datasets without massive manual correction. While not ready for high-precision scholarly work, this level of accuracy can significantly accelerate AI training workflows and improve downstream model performance.

What to watch next

FineBooks will likely keep refining OCR performance and cost-efficiency, potentially moving closer to scholarly-grade accuracy. Monitoring the adoption of dots.mocr or similar models could be important for teams managing large volumes of scanned historical text. The project’s open-source approach also means builders should watch for integration into popular NLP pipelines and data cleaning tools aimed at AI builders and researchers.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.