paper-with-me

Papers

The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection

2026-01-16 · Wei Ai, Yilong Tan, Yuntao Shou, Tao Meng, Haowen Chen, Zhixiong He, Keqin Li arxiv

In recent years, the rapid evolution of large vision-language models (LVLMs) has driven a paradigm shift in multimodal fake news detection (MFND), transforming it from traditional feature-engineering approaches to unified, end-to-end multimodal reasoning frameworks. Early methods primarily relied on shallow fusion techniques to capture correlations between text and images, but they struggled with high-level semantic understanding and complex cross-modal interactions. The emergence of LVLMs has fundamentally changed this landscape by enabling joint modeling of vision and language with powerful representation learning, thereby enhancing the ability to detect misinformation that leverages both textual narratives and visual content. Despite these advances, the field lacks a systematic survey that traces this transition and consolidates recent developments. To address this gap, this paper provides a comprehensive review of MFND through the lens of LVLMs. We first present a historical perspective, mapping the evolution from conventional multimodal detection pipelines to foundation model-driven paradigms. Next, we establish a structured taxonomy covering model architectures, datasets, and performance benchmarks. Furthermore, we analyze the remaining technical challenges, including interpretability, temporal reasoning, and domain generalization. Finally, we outline future research directions to guide the next stage of this paradigm shift. To the best of our knowledge, this is the first comprehensive survey to systematically document and analyze the transformative role of LVLMs in combating multimodal fake news. The summary of existing methods mentioned is in our Github: \href{https://github.com/Tan-YiLong/Overview-of-Fake-News-Detection}{https://github.com/Tan-YiLong/Overview-of-Fake-News-Detection}.

📄 PDF Abstract BibTeX arXiv:2601.15316

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningDomain GeneralizationMultimodal ReasoningFake News Detection

Similar Papers 제목 키워드 기반

Pure Vision Language Action (VLA) Models: A Comprehensive Survey

2025-09-23 · Dapeng Zhang, Jing Sun, Chenghui Hu, Xiaoyan Wu 외 arxiv

The emergence of Vision Language Action (VLA) models marks a paradigm shift from traditional policy-based control to generalized robotics, reframing Vision Language Models (VLMs) from passive sequence generators into act…

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective

2025-02-03 · Xiaorui Ma, Haoran Xie, S. Joe Qin

The integration of vision-language modalities has been a significant focus in multimodal learning, traditionally relying on Vision-Language Pretrained Models. However, with the advent of Large Language Models (LLMs), the…

Automatic Extraction of Causal Relations from Natural Language Texts: A Comprehensive Survey

2016-05-25 · Nabiha Asghar

Automatic extraction of cause-effect relationships from natural language texts is a challenging open problem in Artificial Intelligence. Most of the early attempts at its solution used manually constructed linguistic and…

Relation Extraction

A Survey on Deep Learning Methods for Robot Vision

2018-03-28 · Javier Ruiz-del-Solar, Patricio Loncomilla, Naiomi Soto

Deep learning has allowed a paradigm shift in pattern recognition, from using hand-crafted features together with statistical classifiers to using general-purpose learning procedures for learning data-driven representati…

Deep LearningSurvey

A Survey of Generative Search and Recommendation in the Era of Large Language Models

2024-04-25 · Yongqi Li, Xinyu Lin, Wenjie Wang, Fuli Feng 외

With the information explosion on the Web, search and recommendation are foundational infrastructures to satisfying users' information needs. As the two sides of the same coin, both revolve around the same core research …