paper-with-me

홈 › Papers

Institutional Books 1.0: A 242B token dataset from Harvard Library's collections, refined for accuracy and usability

2025-06-10 · Matteo Cargnelutti, Catherine Brobston, John Hess, Jack Cushman, Kristi Mukk, Aristana Scourtas, Kyle Courtney, Greg Leppert, Amanda Watson, Martha Whitehead, Jonathan Zittrain

Large language models (LLMs) use data to learn about the world in order to produce meaningful correlations and predictions. As such, the nature, scale, quality, and diversity of the datasets used to train these models, or to support their work at inference time, have a direct impact on their quality. The rapid development and adoption of LLMs of varying quality has brought into focus the scarcity of publicly available, high-quality training data and revealed an urgent need to ground the stewardship of these datasets in sustainable practices with clear provenance chains. To that end, this technical report introduces Institutional Books 1.0, a large collection of public domain books originally digitized through Harvard Library's participation in the Google Books project, beginning in 2006. Working with Harvard Library, we extracted, analyzed, and processed these volumes into an extensively-documented dataset of historic texts. This analysis covers the entirety of Harvard Library's collection scanned as part of that project, originally spanning 1,075,899 volumes written in over 250 different languages for a total of approximately 250 billion tokens. As part of this initial release, the OCR-extracted text (original and post-processed) as well as the metadata (bibliographic, source, and generated) of the 983,004 volumes, or 242B tokens, identified as being in the public domain have been made available. This report describes this project's goals and methods as well as the results of the analyses we performed, all in service of making this historical collection more accessible and easier for humans and machines alike to filter, read and use.

📄 PDF Abstract BibTeX arXiv:2506.08300

Code (1)

instdin/institutional-books-1-pipeline 공식 구현 pytorch

Tasks

Optical Character Recognition (OCR)

Methods 이 논문이 사용한 방법론

Library 설명 없음
Focus 설명 없음
Golden Queue Managers 설명 없음

Similar Papers 제목 키워드 기반

Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale

2026-08-19 · David Lowry-Duda, Matteo Cargnelutti, Catherine Brobston, Salwa Ismail 외 arxiv

Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participation in the Google Books Library project…

Institutional Books - Visual Elements: An open-source pipeline for extracting, classifying, deduplicating, and captioning visual elements from digital book collections

2026-08-19 · Jimmy Mendez, Matteo Cargnelutti, David Lowry-Duda, Catherine Brobston 외 arxiv

Historical book collections contain rich visual elements - such as illustrations, photographs, engravings, and decorative art - that are frequently under-explored in large-scale digitization projects. While Optical Chara…

Smart Library: Identifying Books in a Library using Richly Supervised Deep Scene Text Reading

2016-11-22 · Xiao Yang, Dafang He, Wenyi Huang, Zihan Zhou 외

Physical library collections are valuable and long standing resources for knowledge and learning. However, managing books in a large bookshelf and finding books on it often leads to tedious manual work, especially for la…

ManagementRetrievalScene Text DetectionScene Text Recognition+1

LCSHBench: A Multilingual, Consensus-Grounded Benchmark for Library of Congress Subject Heading Assignment

2026-06-03 · Kwok Leong Tang arxiv

Automated subject cataloging assigns controlledvocabulary headings to bibliographic records, but LCSH has no standard public benchmark. We introduce LCSHBench: 22,346 books in 15 languages from the openly licensed Harvar…

Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers

2026-08-19 · Matteo Cargnelutti, Catherine Brobston, Eben English, Jake Sadow 외 arxiv

Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts make computational access to these materials both challenging and limited. We present the Institutional …

Reading Order Detection