paper-with-me

홈 › Papers

Interpretable Visual Understanding with Cognitive Attention Network

2021-08-06 · Xuejiao Tang, Wenbin Zhang, Yi Yu, Kea Turner, Tyler Derr, Mengyu Wang, Eirini Ntoutsi

While image understanding on recognition-level has achieved remarkable advancements, reliable visual scene understanding requires comprehensive image understanding on recognition-level but also cognition-level, which calls for exploiting the multi-source information as well as learning different levels of understanding and extensive commonsense knowledge. In this paper, we propose a novel Cognitive Attention Network (CAN) for visual commonsense reasoning to achieve interpretable visual understanding. Specifically, we first introduce an image-text fusion module to fuse information from images and text collectively. Second, a novel inference module is designed to encode commonsense among image, query and response. Extensive experiments on large-scale Visual Commonsense Reasoning (VCR) benchmark dataset demonstrate the effectiveness of our approach. The implementation is publicly available at https://github.com/tanjatang/CAN

📄 PDF Abstract BibTeX arXiv:2108.02924

Code (1)

tanjatang/CAN 공식 구현 pytorch

Tasks

Scene UnderstandingVisual Commonsense Reasoning

Similar Papers 제목 키워드 기반

Visual Categorization Across Minds and Models: Cognitive Analysis of Human Labeling and Neuro-Symbolic Integration

2025-12-10 · Chethana Prasad Kabgere arxiv

Understanding how humans and AI systems interpret ambiguous visual stimuli offers critical insight into the nature of perception, reasoning, and decision-making. This paper examines image labeling performance across huma…

HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor Generation

2025-11-21 · Jiajun Zhang, Shijia Luo, Ruikang Zhang, Qi Su arxiv

Humor, as both a creative human activity and a social binding mechanism, has long posed a major challenge for AI generation. Although producing humor requires complex cognitive reasoning and social understanding, theorie…

Image CaptioningSemantic Parsing

Shifting Focus with HCEye: Exploring the Dynamics of Visual Highlighting and Cognitive Load on User Attention and Saliency Prediction

2024-04-22 · Anwesha Das, Zekun Wu, Iza Škrjanec, Anna Maria Feit

Visual highlighting can guide user attention in complex interfaces. However, its effectiveness under limited attentional capacities is underexplored. This paper examines the joint impact of visual highlighting (permanent…

Saliency Prediction

Modeling Induced Pleasure through Cognitive Appraisal Prediction via Multimodal Fusion

2026-04-26 · Nastaran Dab, Raziyeh Zall, Mohammadreza Kangavari arxiv

Multimodal affective computing analyzes user-generated social media content to predict emotional states. However, a critical gap remains in understanding how visual content shapes cognitive interpretations and elicits sp…

Attention at Rest Stays at Rest: Breaking Visual Inertia for Cognitive Hallucination Mitigation

2026-04-02 · Boyang Gong, Yu Zheng, Fanye Kong, Jie Zhou 외 arxiv

Like a body at rest that stays at rest, we find that visual attention in multimodal large language models (MLLMs) exhibits pronounced inertia, remaining largely static once settled during early decoding steps and failing…