paper-with-me

Papers

SPARK: Multi-Vision Sensor Perception and Reasoning Benchmark for Large-scale Vision-Language Models

2024-08-22 · Youngjoon Yu, Sangyun Chung, Byung-Kwan Lee, Yong Man Ro

Large-scale Vision-Language Models (LVLMs) have significantly advanced with text-aligned vision inputs. They have made remarkable progress in computer vision tasks by aligning text modality with vision inputs. There are also endeavors to incorporate multi-vision sensors beyond RGB, including thermal, depth, and medical X-ray images. However, we observe that current LVLMs view images taken from multi-vision sensors as if they were in the same RGB domain without considering the physical characteristics of multi-vision sensors. They fail to convey the fundamental multi-vision sensor information from the dataset and the corresponding contextual knowledge properly. Consequently, alignment between the information from the actual physical environment and the text is not achieved correctly, making it difficult to answer complex sensor-related questions that consider the physical environment. In this paper, we aim to establish a multi-vision Sensor Perception And Reasoning benchmarK called SPARK that can reduce the fundamental multi-vision sensor information gap between images and multi-vision sensors. We generated 6,248 vision-language test samples to investigate multi-vision sensory perception and multi-vision sensory reasoning on physical sensor knowledge proficiency across different formats, covering different types of sensor-related questions. We utilized these samples to assess ten leading LVLMs. The results showed that most models displayed deficiencies in multi-vision sensory reasoning to varying extents. Codes and data are available at https://github.com/top-yun/SPARK

📄 PDF Abstract BibTeX arXiv:2408.12114

Code (1)

top-yun/spark 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Are Vision-Language Models Truly Understanding Multi-vision Sensor?

2024-12-30 · Sangyun Chung, Youngjoon Yu, Youngchae Chee, Se Yeon Kim 외

Large-scale Vision-Language Models (VLMs) have advanced by aligning vision inputs with text, significantly improving performance in computer vision tasks. Moreover, for VLMs to be effectively utilized in real-world appli…

Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs

2024-06-20 · Yuxuan Qiao, Haodong Duan, Xinyu Fang, Junming Yang 외

Vision Language Models (VLMs) demonstrate remarkable proficiency in addressing a wide array of visual questions, which requires strong perception and reasoning faculties. Assessing these two competencies independently is…

Language ModellingLarge Language Model

Enhancing Advanced Visual Reasoning Ability of Large Language Models

2024-09-21 · Zhiyuan Li, Dongnan Liu, Chaoyi Zhang, Heng Wang 외

Recent advancements in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models' advanced reasoning ability. Traditional Vision-Language Models (VLMs) perform well in vis…

In-Context LearningVisual Reasoning

Unbiased Visual Reasoning with Controlled Visual Inputs

2025-12-19 · Zhaonan Li, Shijie Lu, Fei Wang, Jacob Dineen 외 arxiv

End-to-end Vision-language Models (VLMs) often answer visual questions by exploiting spurious correlations instead of causal visual evidence, and can become more shortcut-prone when fine-tuned. We introduce VISTA (Visual…

Reinforcement LearningVisual Reasoning

Unlocking Cognitive Capabilities and Analyzing the Perception-Logic Trade-off

2026-02-27 · Longyin Zhang, Shuo Sun, Yingxu He, Won Cheng Yi Lewis 외 arxiv

Recent advancements in Multimodal Large Language Models (MLLMs) pursue omni-perception capabilities, yet integrating robust sensory grounding with complex reasoning remains a challenge, particularly for underrepresented …