paper-with-me

Papers

Video Language Co-Attention with Multimodal Fast-Learning Feature Fusion for VideoQA

2022-05-01 · RepL4NLP (ACL) 2022 5 · Adnen Abdessaied, Ekta Sood, Andreas Bulling

We propose the Video Language Co-Attention Network (VLCN) – a novel memory-enhanced model for Video Question Answering (VideoQA). Our model combines two original contributions”:" A multi-modal fast-learning feature fusion (FLF) block and a mechanism that uses self-attended language features to separately guide neural attention on both static and dynamic visual features extracted from individual video frames and short video clips. When trained from scratch, VLCN achieves competitive results with the state of the art on both MSVD-QA and MSRVTT-QA with 38.06% and 36.01% test accuracies, respectively. Through an ablation study, we further show that FLF improves generalization across different VideoQA datasets and performance for question types that are notoriously challenging in current datasets, such as long questions that require deeper reasoning as well as questions with rare answers.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVideo Question Answering

Similar Papers 제목 키워드 기반

Unlocking Financial Insights: An advanced Multimodal Summarization with Multimodal Output Framework for Financial Advisory Videos

2025-09-25 · Sarmistha Das, R E Zera Marveen Lyngkhoi, Sriparna Saha, Alka Maurya arxiv

The dynamic propagation of social media has broadened the reach of financial advisory content through podcast videos, yet extracting insights from lengthy, multimodal segments (30-40 minutes) remains challenging. We intr…

Speaker Diarization

Attention-Based Multimodal Fusion for Video Description

2017-01-11 · ICCV 2017 10 · Chiori Hori, Takaaki Hori, Teng-Yok Lee, Kazuhiro Sumi 외

Currently successful methods for video description are based on encoder-decoder sentence generation using recur-rent neural networks (RNNs). Recent work has shown the advantage of integrating temporal and/or spatial atte…

DecoderSentenceVideo Description

MicroEmo: Time-Sensitive Multimodal Emotion Recognition with Micro-Expression Dynamics in Video Dialogues

2024-07-23 · Liyun Zhang

Multimodal Large Language Models (MLLMs) have demonstrated remarkable multimodal emotion recognition capabilities, integrating multimodal cues from visual, acoustic, and linguistic contexts in the video to recognize huma…

Emotion RecognitionMultimodal Emotion Recognition

FV2ES: A Fully End2End Multimodal System for Fast Yet Effective Video Emotion Recognition Inference

2022-09-21 · Qinglan Wei, Xuling Huang, Yuan Zhang

In the latest social networks, more and more people prefer to express their emotions in videos through text, speech, and rich facial expressions. Multimodal video emotion analysis techniques can help understand users' in…

Emotion RecognitionMultimodal Emotion RecognitionVideo Emotion Recognition

Mobile-VideoGPT: Fast and Accurate Video Understanding Language Model

2025-03-27 · Abdelrahman Shaker, Muhammad Maaz, Chenhui Gou, Hamid Rezatofighi 외

Video understanding models often struggle with high computational requirements, extensive parameter counts, and slow inference speed, making them inefficient for practical use. To tackle these challenges, we propose Mobi…

EgoSchemaLanguage ModelingLanguage ModellingMVBench+2