paper-with-me

Papers Text Retrieval

“Text Retrieval” 태그가 달린 논문 788편 · 필터 해제

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

2026-08-27 · Chanho Park, Daehyeon Choi, Jihyun Lee, Minhyuk Sung arxiv

Vision-language models (VLMs) can locate an image region referred to by a text prompt and route the corresponding visual evidence to the output, yet the internal mechanism behind this behavior is not understood. Inspired…

Text Retrieval

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

2026-08-26 · Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao, Hiromi Wakaki 외 arxiv

Pretrained vision-language models such as CLIP excel at zero-shot recognition but often fail at compositionality, particularly attribute-object and relational structures. Recent studies mitigate this issue by augmenting …

Text Retrieval

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

2026-08-25 · Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic 외 hf

We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 …

Text Retrieval

ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval

2026-08-16 · Peng Chunyi, Xu Zhipeng, Yan Yukun, Liu Zhenghao 외 hf

Visual document retrieval is a critical component of multimodal retrieval-augmented generation, aiming to identify query-relevant pages from document collections where evidence is distributed across text, layout, charts,…

Representation LearningText Retrieval

DARAD: Dual Adapters and Ranking-Aware Distillation for Continual Remote Sensing Image-Text Retrieval

2026-08-06 · Xi Chen, Xu Chen, Xiangyang Jia, Wei Wang 외 arxiv

With the rapid growth of Earth observation technologies, remote sensing archives are rapidly expanding, making remote sensing image-text retrieval (RS-ITR) increasingly important. However, continual RS-ITR remains challe…

Continual LearningText Retrieval

Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations

2026-08-01 · Shalom Kachko, Raz Lapid, Margarita Vald, Almog Dubin 외 arxiv

Vision-language models (VLMs) process image patches and text tokens in a shared residual stream, but the local geometry through which the two modalities interact remains poorly understood. Most interpretability methods i…

Text Retrieval

A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

2026-07-30 · Jia Yu, Yan Zhu, Yili He, Zilong Wang 외 arxiv

Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather …

Text Retrieval

When Does Explicit View Routing Work? A Controlled Study of Multi-View Graph-Text Alignment

2026-07-29 · Xiao Yue, Guangzhi Qu arxiv

Graph-text retrieval typically maps a graph and its description to a single embedding, even when a query concerns only one semantic aspect, such as a class label or molecular property. Multiple heads can separate these a…

Text Retrieval

Dataset Distillation by Influence Matching

2026-07-18 · Haoru Tan, Wang Wang, Sitong Wu, Xiuzhe Wu 외 hf

We revisit dataset distillation from an outcome-centric perspective. Rather than aligning process surrogates (per-step gradients or training trajectories), Influence Matching (Inf-Match) aligns the final outcome of train…

Text Retrieval

AlphaWiSE: Adaptive Weight Interpolation for Continual Multimodal Representation Learning

2026-07-16 · Sarthak Jain, Qiran Hu, Zhen Zhu, Yaoyao Liu arxiv

Multimodal models such as CLIP learn a shared embedding space for cross-modal retrieval, but continual adaptation to sequentially arriving data can disrupt the cross-modal alignment acquired from earlier phases. Conventi…

Representation LearningCross-Modal RetrievalText Retrieval

What Images Cannot Say: Language-Guided Olfactory Representation Learning

2026-07-07 · Eleftherios Tsonis, Xi Wang, Vicky Kalogeiton arxiv

Images tell us what a scene looks like, but rarely what it would feel like to be there. While recent datasets pair visual scenes with electronic-nose measurements, aligning smell signals with images remains challenging b…

Representation LearningText Retrieval

Submitted and Diagnostic Analysis of Full-Text Temporal Retrieval for LongEval-Sci

2026-07-05 · Yingdong Yang, Haijian Wu arxiv

LongEval-Sci evaluates scientific retrieval under collection change, where a system should be effective on the current corpus and remain usable as documents accumulate over time. This paper reports both official Task 1 r…

Text Retrieval

PLACEMEM: Toward a Compute-Aware Memory Plane for Lifelong Agents

2026-07-05 · Sukanta Ganguly arxiv

Lifelong agents need more than larger context windows and better retrieval. They need memories that can persist, evolve, and be corrected without forcing the serving stack to recompute the same history on every turn or s…

Text Retrieval

HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models

2026-06-22 · Hoang-Bao Le, Aiden Durrant, Thai Son Mai, Binh T. Nguyen 외 arxiv

Vision-Language Models (VLMs) are typically pre-trained on large-scale image-text datasets to capture semantic correspondences between visual content and natural language. However, they remain surprisingly brittle to neg…

Text Retrieval

FusionRS: A Large-Scale RGB-Infrared Remote Sensing Dataset for Dual-Modal Vision-Language Foundation Models

2026-06-15 · Jiaju Han, Ben Zhang, Xuemeng Sun, Qike Zhang 외 arxiv

Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains centered on RGB imagery, leaving the complementary information in infrared data underexplored. Infrared …

Representation LearningText Retrieval

MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA

2026-06-14 · Yilong Zuo, Xunkai Li, Jing Yuan, Qiangqiang Dai 외 arxiv

Long-document multimodal question answering requires a system to locate sparse evidence in long PDFs and integrate clues from text, tables, images, charts, and complex layouts. Existing RAG methods mostly rely on fixed T…

Question AnsweringText Retrieval

Beyond Vector Similarity: A Structural Analysis of Graph-Augmented Retrieval for Industrial Knowledge Graphs

2026-06-04 · Grama Chethan arxiv

Retrieval-Augmented Generation (RAG) fails systematically on queries requiring structural reasoning over interconnected entities. We compare eight retrieval architectures for aerospace supply chain intelligence, progress…

Knowledge GraphsText Retrieval

TechRAG: Evidence-Gated Multimodal Agentic RAG for Technical Literature Reasoning

2026-06-01 · Kanwar Bharat Singh arxiv

This paper presents an agentic multimodal retrieval-augmented generation (RAG) framework for domain-specific literature reasoning, instantiated on a curated corpus of several thousand papers in intelligent tires, vehicle…

Text Retrieval

Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures

2026-05-28 · Varsha Suresh, Mohammad Mahdi Abootorabi, Mohamed Salman, M. Hamza Mughal 외 arxiv

Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains challenging for semantically meaningful gestures whose communicative i…

Gesture GenerationText Retrieval

Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models

2026-05-26 · Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao, Yiming Du 외 arxiv

Large vision-language models increasingly rely on long-context modeling to reason over documents, hour-level videos, and long-horizon agent trajectories, requiring them to locate relevant evidence across interleaved text…

Image RetrievalHead DetectionText Retrieval
1–20 / 788 다음 →