paper-with-me

홈 › Papers

An Automatic Approach for Generating Rich, Linked Geo-Metadata from Historical Map Images

2021-12-03 · Zekun Li, Yao-Yi Chiang, Sasan Tavakkol, Basel Shbita, Johannes H. Uhl, Stefan Leyk, Craig A. Knoblock

Historical maps contain detailed geographic information difficult to find elsewhere covering long-periods of time (e.g., 125 years for the historical topographic maps in the US). However, these maps typically exist as scanned images without searchable metadata. Existing approaches making historical maps searchable rely on tedious manual work (including crowd-sourcing) to generate the metadata (e.g., geolocations and keywords). Optical character recognition (OCR) software could alleviate the required manual work, but the recognition results are individual words instead of location phrases (e.g., "Black" and "Mountain" vs. "Black Mountain"). This paper presents an end-to-end approach to address the real-world problem of finding and indexing historical map images. This approach automatically processes historical map images to extract their text content and generates a set of metadata that is linked to large external geospatial knowledge bases. The linked metadata in the RDF (Resource Description Framework) format support complex queries for finding and indexing historical maps, such as retrieving all historical maps covering mountain peaks higher than 1,000 meters in California. We have implemented the approach in a system called mapKurator. We have evaluated mapKurator using historical maps from several sources with various map styles, scales, and coverage. Our results show significant improvement over the state-of-the-art methods. The code has been made publicly available as modules of the Kartta Labs project at https://github.com/kartta-labs/Project.

📄 PDF Abstract BibTeX arXiv:2112.01671

Code (1)

kartta-labs/project 공식 구현

Tasks

Optical Character RecognitionOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

Towards a Comprehensive Assessment of the Quality and Richness of the Europeana Metadata of food-related Images

2020-05-01 · LREC 2020 5 · Yalemisew Abgaz, Amelie Dorn, Jose Luis Preza Diaz, Gerda Koch

Semantic enrichment of historical images to build interactive AI systems for the Digital Humanities domain has recently gained significant attention. However, before implementing any semantic enrichment tool for building…

WeDH - a Friendly Tool for Building Literary Corpora Enriched with Encyclopedic Metadata

2020-05-01 · LREC 2020 5 · Mattia Egloff, Davide Picca

In recent years the interest in the use of repositories of literary works has been successful. While many efforts related to Linked Open Data go in the right direction, the use of these repositories for the creation of t…

Recommending Scientific Videos based on Metadata Enrichment using Linked Open Data

2018-06-19 · Justyna Medrek, Christian Otto, Ralph Ewerth

The amount of available videos in the Web has significantly increased not only for entertainment etc., but also to convey educational or scientific information in an effective way. There are several web portals that offe…

Optical Character RecognitionOptical Character Recognition (OCR)speech-recognitionSpeech Recognition

Linking Hadith Narrator Identities Across Heterogeneous Arabic Biographical Databases: A Multi-Signal Entity Resolution Pipeline

2026-06-30 · Taufiq Wirahman arxiv

The transmission chains (sanad) of Islamic Hadith literature encode relationships among tens of thousands of historical narrators whose biographical records are dispersed across independently maintained digital databases…

Entity Resolution

KwicKwocKwac, a tool for rapidly generating concordances and marking up a literary text

2024-10-08 · Sebastian Barzaghi, Francesco Paolucci, Francesca Tomasi, Fabio Vitali

This paper introduces KwicKwocKwac 1.0 (KwicKK), a web application designed to enhance the annotation and enrichment of digital texts in the humanities. KwicKK provides a user-friendly interface that enables scholars and…