paper-with-me

홈 › Papers

Integrating Visual and Textual Inputs for Searching Large-Scale Map Collections with CLIP

2024-10-02 · Jamie Mahowald, Benjamin Charles Germain Lee

Despite the prevalence and historical importance of maps in digital collections, current methods of navigating and exploring map collections are largely restricted to catalog records and structured metadata. In this paper, we explore the potential for interactively searching large-scale map collections using natural language inputs ("maps with sea monsters"), visual inputs (i.e., reverse image search), and multimodal inputs (an example map + "more grayscale"). As a case study, we adopt 562,842 images of maps publicly accessible via the Library of Congress's API. To accomplish this, we use the mulitmodal Contrastive Language-Image Pre-training (CLIP) machine learning model to generate embeddings for these maps, and we develop code to implement exploratory search capabilities with these input strategies. We present results for example searches created in consultation with staff in the Library of Congress's Geography and Map Division and describe the strengths, weaknesses, and possibilities for these search queries. Moreover, we introduce a fine-tuning dataset of 10,504 map-caption pairs, along with an architecture for fine-tuning a CLIP model on this dataset. To facilitate re-use, we provide all of our code in documented, interactive Jupyter notebooks and place all code into the public domain. Lastly, we discuss the opportunities and challenges for applying these approaches across both digitized and born-digital collections held by galleries, libraries, archives, and museums.

📄 PDF Abstract BibTeX arXiv:2410.01190

Code (1)

j-mahowald/clip-loc-maps 공식 구현 pytorch

Tasks

Image Retrieval

Methods 이 논문이 사용한 방법론

Library 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports From Scratch with Agentic Framework

2025-06-03 · Zhaorui Yang, Bo Pan, Han Wang, Yiyao Wang 외

Visualizations play a crucial part in effective communication of concepts and information. Recent advances in reasoning and retrieval augmented generation have enabled Large Language Models (LLMs) to perform deep researc…

Retrieval-augmented Generation

The Contemporary Art of Image Search: Iterative User Intent Expansion via Vision-Language Model

2023-12-04 · Yilin Ye, Qian Zhu, Shishi Xiao, Kang Zhang 외

Image search is an essential and user-friendly method to explore vast galleries of digital images. However, existing image search methods heavily rely on proximity measurements like tag matching or image similarity, requ…

Image RetrievalInteractive SegmentationLanguage ModelingLanguage Modelling+1

Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning

2024-12-15 · Shengqiong Wu, Hao Fei, Liangming Pan, William Yang Wang 외

Recent advancements in multimodal large language models (MLLMs) have shown unprecedented capabilities in advancing various vision-language tasks. However, MLLMs face significant challenges with hallucinations, and mislea…

Hallucination

Let Androids Dream of Electric Sheep: A Human-like Image Implication Understanding and Reasoning Framework

2025-05-22 · Chenhao Zhang, Yazhe Niu

Metaphorical comprehension in images remains a critical challenge for AI systems, as existing models struggle to grasp the nuanced cultural, emotional, and contextual implications embedded in visual content. While multim…

Multiple-choiceVisual Question Answering (VQA)

Rethinking Homogeneity of Vision and Text Tokens in Large Vision-and-Language Models

2025-02-04 · Chia-Wen Kuo, Sijie Zhu, Fan Chen, Xiaohui Shen 외

Large vision-and-language models (LVLMs) typically treat visual and textual embeddings as homogeneous inputs to a large language model (LLM). However, these inputs are inherently different: visual inputs are multi-dimens…

Language ModelingLanguage ModellingLarge Language Model