paper-with-me

홈 › Papers

Allegory of the Cave: Measurement-Grounded Vision-Language Learning

2026-05-12 · Kepeng Xu, Li Xu, Gang He, Wenxin Yu arxiv

Vision-language models typically reason over post-ISP RGB images, although RGB rendering can clip, suppress, or quantize sensor evidence before inference. We study whether grounding improves when the visual interface is moved closer to the underlying camera measurement. We formulate measurement-grounded vision-language learning and instantiate it as PRISM-VL, which combines RAW-derived Meas.-XYZ inputs, camera-conditioned grounding, and Exposure-Bracketed Supervision Aggregation for transferring supervision from RGB proxies to measurement-domain observations. Using a quality-controlled 150K instruction-tuning set and a held-out benchmark targeting low-light, HDR, visibility-sensitive, and hallucination-sensitive cases, PRISM-VL-8B reaches 0.6120 BLEU, 0.4571 ROUGE-L, and 82.66\% LLM-Judge accuracy, improving over the RGB Qwen3-VL-8B baseline by +0.1074 BLEU, +0.1071 ROUGE-L, and +4.46 percentage points. These results suggest that part of VLM grounding error arises from information lost during RGB rendering, and that preserving measurement-domain evidence can improve multimodal reasoning.

📄 PDF Abstract BibTeX arXiv:2605.11727

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework

2026-06-16 · Sneha Rao, Shaina Raza, Dhanesh Ramachandram arxiv

Vision-Language Models (VLMs) remain prone to hallucinations, producing fluent but visually unfaithful outputs. Existing chain-of-thought and retrieval-augmented methods only partially address this, as they neither enfor…

Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts

2025-01-02 · Youcheng Huang, Chen Huang, Duanyu Feng, Wenqiang Lei 외

Understanding the inner workings of Large Language Models (LLMs) is a critical research frontier. Prior research has shown that a single LLM's concept representations can be captured as steering vectors (SVs), enabling t…

CAVE-NAV: VLM-Based Autonomous 3D Navigation in Underwater Cave Environments

2026-08-28 · Zhenqi Wu, Yuanjie Lu, Yisheng Zhang, Miao Yu 외 arxiv

Autonomous navigation in underwater cave environments is essential for search-and-rescue operations, scientific exploration, and emergency egress. Traditional navigation systems commonly depend on dense visual features f…

CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments

2025-10-29 · Rishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges 외 arxiv

Humans can naturally identify, reason about, and explain anomalies in their environment. In computer vision, this long-standing challenge remains limited to industrial defects or unrealistic, synthetically generated anom…

Anomaly DetectionVisual Grounding

EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography

2025-12-13 · Yuheng Li, Yue Zhang, Abdoul Aziz Amadou, Yuxiang Lai 외 arxiv

Echocardiography is the most widely used imaging modality in cardiology, yet its interpretation remains labor-intensive and inherently multimodal, requiring view recognition, quantitative measurements, qualitative assess…

Text Retrieval