paper-with-me

홈 › Papers

Are Vision-Language Models Truly Understanding Multi-vision Sensor?

2024-12-30 · Sangyun Chung, Youngjoon Yu, Youngchae Chee, Se Yeon Kim, Byung-Kwan Lee, Yong Man Ro

Large-scale Vision-Language Models (VLMs) have advanced by aligning vision inputs with text, significantly improving performance in computer vision tasks. Moreover, for VLMs to be effectively utilized in real-world applications, an understanding of diverse multi-vision sensor data, such as thermal, depth, and X-ray information, is essential. However, we find that current VLMs process multi-vision sensor images without deep understanding of sensor information, disregarding each sensor's unique physical properties. This limitation restricts their capacity to interpret and respond to complex questions requiring multi-vision sensor reasoning. To address this, we propose a novel Multi-vision Sensor Perception and Reasoning (MS-PR) benchmark, assessing VLMs on their capacity for sensor-specific reasoning. Moreover, we introduce Diverse Negative Attributes (DNA) optimization to enable VLMs to perform deep reasoning on multi-vision sensor tasks, helping to bridge the core information gap between images and sensor data. Extensive experimental results validate that the proposed DNA method can significantly improve the multi-vision sensor reasoning for VLMs.

📄 PDF Abstract BibTeX arXiv:2412.20750

Code (1)

top-yun/ms-pr 공식 구현 pytorch

Similar Papers 제목 키워드 기반

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

2026-06-24 · Haoxiang Sun, Tao Wang, Li Yuan, Jian Zhao 외 arxiv

Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of models such as OpenAI's O-series and DeepS…

Weakly Supervised POS Taggers Perform Poorly on Truly Low-Resource Languages

2020-04-28 · Katharina Kann, Ophélie Lacroix, Anders Søgaard

Part-of-speech (POS) taggers for low-resource languages which are exclusively based on various forms of weak supervision - e.g., cross-lingual transfer, type-level supervision, or a combination thereof - have been report…

Cross-Lingual TransferPOSPOS Tagging

Left-Right Symmetry Breaking in CLIP-style Vision-Language Models Trained on Synthetic Spatial-Relation Data

2026-01-19 · Takaki Yamamoto, Chihiro Noguchi, Toshihiro Tanizawa arxiv

Spatial understanding remains a key challenge in vision-language models. Yet it is still unclear whether such understanding is truly acquired, and if so, through what mechanisms. We present a controllable 1D image-text t…

Unlocking Dense Metric Depth Estimation in VLMs

2026-05-15 · Hanxun Yu, Xuan Qu, Yuxin Wang, Jianke Zhu 외 arxiv

Vision-Language Models (VLMs) excel at 2D tasks such as grounding and captioning, yet remain limited in 3D understanding. A key limitation is their text-only supervision paradigm, which under-constrains fine-grained visu…

Spatial ReasoningDepth Estimation

Don't Buy it! Reassessing the Ad Understanding Abilities of Contrastive Multimodal Models

2024-05-31 · A. Bavaresco, A. Testoni, R. Fernández

Image-based advertisements are complex multimodal stimuli that often contain unusual visual elements and figurative language. Previous research on automatic ad understanding has reported impressive zero-shot accuracy of …

Multimodal ReasoningRetrieval