paper-with-me

홈 › Papers

Seeing, Hearing, and Knowing Together: Multimodal Strategies in Deepfake Videos Detection

2026-02-01 · Chen Chen, Dion Hoe-Lian Goh arxiv

As deepfake videos become increasingly difficult for people to recognise, understanding the strategies humans use is key to designing effective media literacy interventions. We conducted a study with 195 participants between the ages of 21 and 40, who judged real and deepfake videos, rated their confidence, and reported the cues they relied on across visual, audio, and knowledge strategies. Participants were more accurate with real videos than with deepfakes and showed lower expected calibration error for real content. Through association rule mining, we identified cue combinations that shaped performance. Visual appearance, vocal, and intuition often co-occurred for successful identifications, which highlights the importance of multimodal approaches in human detection. Our findings show which cues help or hinder detection and suggest directions for designing media literacy tools that guide effective cue use. Building on these insights can help people improve their identification skills and become more resilient to deceptive digital media.

📄 PDF Abstract BibTeX arXiv:2602.01284

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners

2024-02-27 · CVPR 2024 1 · Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang 외

Video and audio content creation serves as the core technique for the movie industry and professional users. Recently, existing diffusion-based methods tackle video and audio generation separately, which hinders the tech…

Audio GenerationDenoising

Seeing and Hearing Egocentric Actions: How Much Can We Learn?

2019-10-15 · Alejandro Cartas, Jordi Luque, Petia Radeva, Carlos Segura 외

Our interaction with the world is an inherently multimodal experience. However, the understanding of human-to-object interactions has historically been addressed focusing on a single modality. In particular, a limited nu…

Action Recognition

Seeing Sound, Hearing Sight: Uncovering Modality Bias and Conflict of AI models in Sound Localization

2025-05-16 · Yanhao Jia, Ji Xie, S Jivaganesh, Hao Li 외

Imagine hearing a dog bark and turning toward the sound only to see a parked car, while the real, silent dog sits elsewhere. Such sensory conflicts test perception, yet humans reliably resolve them by prioritizing sound …

Seeing and hearing what has not been said; A multimodal client behavior classifier in Motivational Interviewing with interpretable fusion

2023-09-25 · Lucie Galland, Catherine Pelachaud, Florian Pecune

Motivational Interviewing (MI) is an approach to therapy that emphasizes collaboration and encourages behavioral change. To evaluate the quality of an MI conversation, client utterances can be classified using the MISC c…

Decision Making

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

2026-07-28 · Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan 외 arxiv

Multimodal Large Language Models (MLLMs) achieve strong performance by integrating visual inputs with the rich priors of pretrained language models. However, they often fail on vision-centric tasks, especially when visua…

Image Reconstruction