paper-with-me

홈 › Papers

Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models

2025-11-24 · Jonathan Lee, Xingrui Wang, Jiawei Peng, Luoxin Ye, Zehan Zheng, Tiezheng Zhang, Tao Wang, Wufei Ma, Siyi Chen, Yu-Cheng Chou, Prakhar Kaushik, Alan Yuille arxiv

We propose Perceptual Taxonomy, a structured process of scene understanding that first recognizes objects and their spatial configurations, then infers task-relevant properties such as material, affordance, function, and physical attributes to support goal-directed reasoning. While this form of reasoning is fundamental to human cognition, current vision-language benchmarks lack comprehensive evaluation of this ability and instead focus on surface-level recognition or image-text alignment. To address this gap, we introduce Perceptual Taxonomy, a benchmark for physically grounded visual reasoning. We annotate 3173 objects with four property families covering 84 fine-grained attributes. Using these annotations, we construct a multiple-choice question benchmark with 5802 images across both synthetic and real domains. The benchmark contains 28033 template-based questions spanning four types (object description, spatial reasoning, property matching, and taxonomy reasoning), along with 50 expert-crafted questions designed to evaluate models across the full spectrum of perceptual taxonomy reasoning. Experimental results show that leading vision-language models perform well on recognition tasks but degrade by 10 to 20 percent on property-driven questions, especially those requiring multi-step reasoning over structured attributes. These findings highlight a persistent gap in structured visual understanding and the limitations of current models that rely heavily on pattern matching. We also show that providing in-context reasoning examples from simulated scenes improves performance on real-world and expert-curated questions, demonstrating the effectiveness of perceptual-taxonomy-guided prompting.

📄 PDF Abstract BibTeX arXiv:2511.19526

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingSpatial ReasoningVisual Reasoning

Similar Papers 제목 키워드 기반

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

2026-05-18 · Yuqi Tang, Yang Shi, Zhuoran Zhang, Qixun Wang 외 arxiv

Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inconsistencies, structural distortions, and semantic incoherence. While…

Video ClassificationArtifact Detection

Photon-Driven Neural Path Guiding

2020-10-05 · Shilin Zhu, Zexiang Xu, Tiancheng Sun, Alexandr Kuznetsov 외

Although Monte Carlo path tracing is a simple and effective algorithm to synthesize photo-realistic images, it is often very slow to converge to noise-free results when involving complex global illumination. One of the m…

Hierarchical learning for DNN-based acoustic scene classification

2016-07-13 · Yong Xu, Qiang Huang, Wenwu Wang, Mark D. Plumbley

In this paper, we present a deep neural network (DNN)-based acoustic scene classification framework. Two hierarchical learning methods are proposed to improve the DNN baseline performance by incorporating the hierarchica…

Acoustic Scene ClassificationClassificationGeneral ClassificationScene Classification

A New Method for Evaluating Automatically Learned Terminological Taxonomies

2012-05-01 · LREC 2012 5 · Paola Velardi, Roberto Navigli, Stefano Faralli, Juana Maria Ruiz Martinez

Abstract Evaluating a taxonomy learned automatically against an existing gold standard is a very complex problem, because differences stem from the number, label, depth and ordering of the taxonomy nodes. In this paper w…

Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning

2025-11-27 · Zhaoyang Wei, Wenchao Ding, Yanchao Hao, Xi Chen arxiv

Models capable of "thinking with images" by dynamically grounding their reasoning in visual evidence represent a major leap in multimodal AI. However, replicating and advancing this ability is non-trivial, with current m…

Reinforcement LearningVisual Reasoning