paper-with-me

Papers

Is Cognition consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding

2024-11-12 · Zirui Shao, Chuwei Luo, Zhaoqing Zhu, Hangdi Xing, Zhi Yu, Qi Zheng, Jiajun Bu

Multimodal large language models (MLLMs) have shown impressive capabilities in document understanding, a rapidly growing research area with significant industrial demand in recent years. As a multimodal task, document understanding requires models to possess both perceptual and cognitive abilities. However, current MLLMs often face conflicts between perception and cognition. Taking a document VQA task (cognition) as an example, an MLLM might generate answers that do not match the corresponding visual content identified by its OCR (perception). This conflict suggests that the MLLM might struggle to establish an intrinsic connection between the information it "sees" and what it "understands." Such conflicts challenge the intuitive notion that cognition is consistent with perception, hindering the performance and explainability of MLLMs. In this paper, we define the conflicts between cognition and perception as Cognition and Perception (C&P) knowledge conflicts, a form of multimodal knowledge conflicts, and systematically assess them with a focus on document understanding. Our analysis reveals that even GPT-4o, a leading MLLM, achieves only 68.6% C&P consistency. To mitigate the C&P knowledge conflicts, we propose a novel method called Multimodal Knowledge Consistency Fine-tuning. This method first ensures task-specific consistency and then connects the cognitive and perceptual knowledge. Our method significantly reduces C&P knowledge conflicts across all tested MLLMs and enhances their performance in both cognitive and perceptual tasks in most scenarios.

📄 PDF Abstract BibTeX arXiv:2411.07722

Code (0)

등록된 구현이 없습니다.

Tasks

document understandingOptical Character Recognition (OCR)Visual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

TiCAL:Typicality-Based Consistency-Aware Learning for Multimodal Emotion Recognition

2025-11-19 · Wen Yin, Siyu Zhan, Cencen Liu, Xin Hu 외 arxiv

Multimodal Emotion Recognition (MER) aims to accurately identify human emotional states by integrating heterogeneous modalities such as visual, auditory, and textual data. Existing approaches predominantly rely on unifie…

Multimodal Emotion Recognition

FaceInsight: A Multimodal Large Language Model for Face Perception

2025-04-22 · Jingzhi Li, Changjiang Luo, Ruoyu Chen, Hua Zhang 외

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, ofte…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

Seeing is Not Understanding: A Benchmark on Perception-Cognition Disparities in Large Language Models

2025-09-14 · Haokun Li, Yazhou Zhang, Jizhi Ding, Qiuchi Li 외 arxiv

With the rapid advancement of Multimodal Large Language Models (MLLMs), they have demonstrated exceptional capabilities across a variety of vision-language tasks. However, current evaluation benchmarks predominantly focu…

Visual Question Answering

XEmoGPT: An Explainable Multimodal Emotion Recognition Framework with Cue-Level Perception and Reasoning

2026-02-05 · Hanwen Zhang, Yao Liu, Peiyuan Jiang, Lang Junjie 외 arxiv

Explainable Multimodal Emotion Recognition plays a crucial role in applications such as human-computer interaction and social media analytics. However, current approaches struggle with cue-level perception and reasoning …

Multimodal Emotion RecognitionSemantic Similarity

From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping

2026-07-01 · Jian Song, Tian Zi, Shen Guanting arxiv

Improvements in the technical performance of human--robot interaction (HRI) systems do not automatically translate into differences that human users can detect during live interaction. This paper investigates whether a 1…

Speech RecognitionObject Detection