paper-with-me

홈 › Papers

Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)

2025-12-22 · Niclas Griesshaber, Jochen Streb arxiv

We leverage multimodal large language models (LLMs) to construct a dataset of 306,070 German patents (1877-1918) from 9,562 archival image scans using our LLM-based pipeline powered by Gemini-2.5-Pro and Gemini-2.5-Flash-Lite. Our benchmarking exercise provides tentative evidence that multimodal LLMs can create higher quality datasets than our research assistants, while also being more than 795 times faster and 205 times cheaper in constructing the patent dataset from our image corpus. About 20 to 50 patent entries are embedded on each page, arranged in a double-column format and printed in Gothic and Roman fonts. The font and layout complexity of our primary source material suggests to us that multimodal LLMs are a paradigm shift in how datasets are constructed in economic history. We open-source our benchmarking and patent datasets as well as our LLM-based data pipeline, which can be easily adapted to other image corpora using LLM-assisted coding tools, lowering the barriers for less technical researchers. Finally, we explain the economics of deploying LLMs for historical dataset construction and conclude by speculating on the potential implications for the field of economic history.

📄 PDF Abstract BibTeX arXiv:2512.19675

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Historical Tabular Image to Knowledge Graphs: A Provenance-Aware Modular Pipeline

2026-05-06 · Sarah Binta Alam Shoilee, Victor de Boer, Jacco van Ossenbruggen, Susan Legêne arxiv

Handwritten archival tables contain rich historical information, yet transforming them into structured representations, such as Knowledge Graphs, requires integrating table structure recognition, handwriting recognition,…

Handwriting RecognitionInformation ExtractionKnowledge Graphs

Coloring the Past: Neural Historical Buildings Reconstruction from Archival Photography

2023-11-29 · David Komorowicz, Lu Sang, Ferdinand Maiwald, Daniel Cremers

Historical buildings are a treasure and milestone of human cultural heritage. Reconstructing the 3D models of these building hold significant value. The rapid development of neural rendering methods makes it possible to …

Neural Rendering

On Path to Multimodal Historical Reasoning: HistBench and HistAgent

2025-05-26 · Jiahao Qiu, Fulian Xiao, Yimin Wang, Yuchen Mao 외

Recent advances in large language models (LLMs) have led to remarkable progress across domains, yet their capabilities in the humanities, particularly history, remain underexplored. Historical reasoning poses unique chal…

Optical Character Recognition (OCR)

PereStruct: Multimodal Semantic Assembly for Robust Historical Document Parsing

2026-06-03 · Maksim Shandybo, Ivan Bespalov, Daniil Yefimov, Marina Kosheleva 외 arxiv

Parsing historical documents with complex, non-standard layouts remains a fundamental bottleneck in large-scale archival digitization. Unlike modern typography, historical newspapers exhibit severe physical degradation a…

Semantic Similarity

WeatherArchive-Bench: Benchmarking Retrieval-Augmented Reasoning for Historical Weather Archives

2025-10-06 · Yongan Yu, Xianda Du, Qingchen Hu, Jiahao Liang 외 arxiv

Historical archives on weather events are collections of enduring primary source records that offer rich, untapped narratives of how societies have experienced and responded to extreme weather events. These qualitative a…