Megadose Built for builders and researchers.

Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale

· ArXiv · AI/CL/LG ·
Harvard’s 983,000-volume Google Books corpus now has an open enriched-text layer for OCR cleanup without forcing one preprocessing policy.

The release covers 983,003 IB-HL volumes, with 217B o200k_base tokens arranged into 1.39B annotated subtopic paragraphs. The pipeline normalizes text while preserving metadata in HTML-like annotations, including endmatter separation, paragraph-level language detection, duplicate paragraph clusters, and bits-per-byte scores. Its authors frame this as an alternative to aggressive web-text preprocessing, letting users decide what to filter or keep. It applies across roughly 250 languages in the collection. ArXiv · AI/CL/LG's note

score 5

Categories: OSS & Tools, Research