paper-with-me

홈 › Papers

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning

2025-05-14 · Dayong Liang, Changmeng Zheng, Zhiyuan Wen, Yi Cai, Xiao-Yong Wei, Qing Li

Traditional scene graphs primarily focus on spatial relationships, limiting vision-language models' (VLMs) ability to reason about complex interactions in visual scenes. This paper addresses two key challenges: (1) conventional detection-to-construction methods produce unfocused, contextually irrelevant relationship sets, and (2) existing approaches fail to form persistent memories for generalizing interaction reasoning to new scenes. We propose Interaction-augmented Scene Graph Reasoning (ISGR), a framework that enhances VLMs' interactional reasoning through three complementary components. First, our dual-stream graph constructor combines SAM-powered spatial relation extraction with interaction-aware captioning to generate functionally salient scene graphs with spatial grounding. Second, we employ targeted interaction queries to activate VLMs' latent knowledge of object functionalities, converting passive recognition into active reasoning about how objects work together. Finally, we introduce a lone-term memory reinforcement learning strategy with a specialized interaction-focused reward function that transforms transient patterns into long-term reasoning heuristics. Extensive experiments demonstrate that our approach significantly outperforms baseline methods on interaction-heavy reasoning benchmarks, with particularly strong improvements on complex scene understanding tasks. The source code can be accessed at https://github.com/open_upon_acceptance.

📄 PDF Abstract BibTeX arXiv:2505.09118

Code (0)

등록된 구현이 없습니다.

Tasks

Relation ExtractionScene Understanding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Seeing things or seeing scenes: Investigating the capabilities of V&L models to align scene descriptions to images

2021-10-16 · ACL ARR October 2021 10 · Anonymous

Images can be described in terms of the objects they contain, or in terms of the types of scene or place that they instantiate. In this paper we address to what extent pretrained Vision and Language models can learn to a…

Object

Seeing Symbols, Missing Cultures: Probing Vision-Language Models' Reasoning on Fire Imagery and Cultural Meaning

2025-09-27 · Haorui Yu, Yang Zhao, Yijia Chu, Qiufeng Yi arxiv

Vision-Language Models (VLMs) often appear culturally competent but rely on superficial pattern matching rather than genuine cultural understanding. We introduce a diagnostic framework to probe VLM reasoning on fire-them…

Seeing Beyond Classes: Zero-Shot Grounded Situation Recognition via Language Explainer

2024-04-24 · JiaMing Lei, Lin Li, Chunping Wang, Jun Xiao 외

Benefiting from strong generalization ability, pre-trained vision language models (VLMs), e.g., CLIP, have been widely utilized in zero-shot scene understanding. Unlike simple recognition tasks, grounded situation recogn…

Grounded Situation RecognitionScene Understanding

Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videos

2026-03-01 · Shreshth Saini, Bowen Chen, Neil Birkbeck, Yilin Wang 외 arxiv

High Dynamic Range (HDR) user-generated (UGC) videos are rapidly proliferating across social platforms, yet most perceptual video quality assessment (VQA) systems remain tailored to Standard Dynamic Range (SDR). HDR has …

Video Quality Assessment

Single Image 3D Without a Single 3D Image

2015-12-01 · ICCV 2015 12 · David F. Fouhey, Wajahat Hussain, Abhinav Gupta, Martial Hebert

Do we really need 3D labels in order to learn how to predict 3D? In this paper, we show that one can learn a mapping from appearance to 3D properties without ever seeing a single explicit 3D label. Rather than use explic…

Scene Understanding