paper-with-me

홈 › Papers

Institutional Books - Visual Elements: An open-source pipeline for extracting, classifying, deduplicating, and captioning visual elements from digital book collections

2026-08-19 · Jimmy Mendez, Matteo Cargnelutti, David Lowry-Duda, Catherine Brobston, Salwa Ismail, Greg Leppert, Amanda Watson, Jonathan Zittrain arxiv

Historical book collections contain rich visual elements - such as illustrations, photographs, engravings, and decorative art - that are frequently under-explored in large-scale digitization projects. While Optical Character Recognition (OCR) has standardized the extraction of textual content, these visual components offer a layer of nuance and context that remains largely untapped by automated text extraction workflows. This technical report introduces Institutional Books - Visual Elements, an open-source end-to-end pipeline for detecting, classifying, deduplicating, and captioning visual elements from historical book collections. Alongside this pipeline, we release an initial dataset of 22.6 million visual elements extracted from the 983,004 scanned volumes that comprise the Institutional Books: Harvard Library dataset. This work contributes to ongoing, community-wide efforts to enable new use cases for digitized library collections through computational access, from artificial intelligence model training to digital humanities research.

📄 PDF Abstract BibTeX arXiv:2608.18957

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Comparing Machine Learning Approaches for Table Recognition in Historical Register Books

2019-06-14 · Stéphane Clinchant, Hervé Déjean, Jean-Luc Meunier, Eva Lang 외

We present in this paper experiments on Table Recognition in hand-written registry books. We first explain how the problem of row and column detection is modeled, and then compare two Machine Learning approaches (Conditi…

BIG-bench Machine LearningTable Recognition

SuperNOVA: Design Strategies and Opportunities for Interactive Visualization in Computational Notebooks

2023-05-04 · Zijie J. Wang, David Munechika, Seongmin Lee, Duen Horng Chau

Computational notebooks, such as Jupyter Notebook, have become data scientists' de facto programming environments. Many visualization researchers and practitioners have developed interactive visualization tools that supp…

Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale

2026-08-19 · David Lowry-Duda, Matteo Cargnelutti, Catherine Brobston, Salwa Ismail 외 arxiv

Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participation in the Google Books Library project…

Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG

2026-05-27 · Yubo Li, Rema Padman, Ramayya Krishnan arxiv

A retrieval-augmented generation (RAG) system deployed over a multi-author institutional corpus can give a different answer to the same question depending on which source it retrieves -- a failure mode the dominant singl…

Knowledge Activation: AI Skills as the Institutional Knowledge Primitive for Agentic Software Development

2026-03-16 · Gal Bakal arxiv

Enterprise software organizations accumulate critical institutional knowledge - architectural decisions, deployment procedures, compliance policies, incident playbooks - yet this knowledge remains trapped in formats desi…