Megadose AI progress, ranked and analyzed.

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers

· HF Daily Papers ·
Boston Public Library scans yielded 16.3 billion OCR tokens from public-domain newspapers.

The paper describes a modular pipeline for turning dense historical newspaper scans into structured, crop-level datasets. It segments pages, runs OCR, then adds classification, reading order, named entities, subjects, language detection, and embeddings. The authors ran it on 1,473,635 scans from Boston Public Library holdings, covering newspapers published from 1795 to 1930. They released the pipeline, models, and resulting open dataset. HF Daily Papers' note

score 4

Categories: Research