paper-with-me

홈 › Papers

Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion

2023-06-15 · Ishaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal, Basura Fernando, Cheston Tan

While VideoQA Transformer models demonstrate competitive performance on standard benchmarks, the reasons behind their success are not fully understood. Do these models capture the rich multimodal structures and dynamics from video and text jointly? Or are they achieving high scores by exploiting biases and spurious features? Hence, to provide insights, we design $\textit{QUAG}$ (QUadrant AveraGe), a lightweight and non-parametric probe, to conduct dataset-model combined representation analysis by impairing modality fusion. We find that the models achieve high performance on many datasets without leveraging multimodal representations. To validate QUAG further, we design $\textit{QUAG-attention}$, a less-expressive replacement of self-attention with restricted token interactions. Models with QUAG-attention achieve similar performance with significantly fewer multiplication operations without any finetuning. Our findings raise doubts about the current models' abilities to learn highly-coupled multimodal representations. Hence, we design the $\textit{CLAVI}$ (Complements in LAnguage and VIdeo) dataset, a stress-test dataset curated by augmenting real-world videos to have high modality coupling. Consistent with the findings of QUAG, we find that most of the models achieve near-trivial performance on CLAVI. This reasserts the limitations of current models for learning highly-coupled multimodal representations, that is not evaluated by the current datasets (project page: https://dissect-videoqa.github.io ).

📄 PDF Abstract BibTeX arXiv:2306.08889

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarkingcounterfactual

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Counterfactuals 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Position-Wise Feed-Forward Layer 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…

Similar Papers 제목 키워드 기반

Rethinking the Potential of Multimodality in Collaborative Problem Solving Diagnosis with Large Language Models

2025-04-21 · K. Wong, B. Wu, S. Bulathwela, M. Cukurova

Detecting collaborative and problem-solving behaviours from digital traces to interpret students' collaborative problem solving (CPS) competency is a long-term goal in the Artificial Intelligence in Education (AIEd) fiel…

Frequency-Aware Vision-Language Multimodality Generalization Network for Remote Sensing Image Classification

2025-11-13 · Junjie Zhang, Feng Zhao, Hanqiang Liu, Jun Yu arxiv

The booming remote sensing (RS) technology is giving rise to a novel multimodality generalization task, which requires the model to overcome data heterogeneity while possessing powerful cross-scene generalization ability…

Remote Sensing Image Classification

Reliable Multimodality Eye Disease Screening via Mixture of Student's t Distributions

2023-03-17 · Ke Zou, Tian Lin, Xuedong Yuan, Haoyu Chen 외

Multimodality eye disease screening is crucial in ophthalmology as it integrates information from diverse sources to complement their respective performances. However, the existing methods are weak in assessing the relia…

Decision Making

Multimodal Motion Prediction with Stacked Transformers

2021-03-22 · CVPR 2021 1 · Yicheng Liu, Jinghuai Zhang, Liangji Fang, Qinhong Jiang 외

Predicting multiple plausible future trajectories of the nearby vehicles is crucial for the safety of autonomous driving. Recent motion prediction approaches attempt to achieve such multimodal motion prediction by implic…

Autonomous DrivingDiversitymotion predictionPrediction

LetsTalk: Latent Diffusion Transformer for Talking Video Synthesis

2024-11-24 · Haojie Zhang, Zhihao Liang, Ruibo Fu, Zhengqi Wen 외

Portrait image animation using audio has rapidly advanced, enabling the creation of increasingly realistic and expressive animated faces. The challenges of this multimodality-guided video generation task involve fusing v…

DiversityImage AnimationVideo Generation