paper-with-me

홈 › Papers

Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting

2024-10-01 · Chen Cai, Zheng Wang, Jianjun Gao, Wenyang Liu, Ye Lu, Runzhong Zhang, Kim-Hui Yap

In recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they struggle to adapt to new questions or tasks posed by newly available content. In this paper, we explore the novel challenge of VideoQA within a continual learning framework, and empirically identify a critical issue: fine-tuning a large language model (LLM) for a sequence of tasks often results in catastrophic forgetting. To address this, we propose Collaborative Prompting (ColPro), which integrates specific question constraint prompting, knowledge acquisition prompting, and visual temporal awareness prompting. These prompts aim to capture textual question context, visual content, and video temporal dynamics in VideoQA, a perspective underexplored in prior research. Experimental results on the NExT-QA and DramaQA datasets show that ColPro achieves superior performance compared to existing approaches, achieving 55.14\% accuracy on NExT-QA and 71.24\% accuracy on DramaQA, highlighting its practical relevance and effectiveness.

📄 PDF Abstract BibTeX arXiv:2410.00771

Code (1)

caicch/colpro 공식 구현 pytorch

Tasks

Continual LearningLanguage ModelingLanguage ModellingLarge Language ModelQuestion AnsweringVideo Question Answering

Similar Papers 제목 키워드 기반

DAM: Dynamic Adapter Merging for Continual Video QA Learning

2024-03-13 · Feng Cheng, Ziyang Wang, Yi-Lin Sung, Yan-Bo Lin 외

We present a parameter-efficient method for continual video question-answering (VidQA) learning. Our method, named DAM, uses the proposed Dynamic Adapter Merging to (i) mitigate catastrophic forgetting, (ii) enable effic…

Continual Learningimage-classificationImage ClassificationQuestion Answering+1

Video Question Answering: Datasets, Algorithms and Challenges

2022-03-02 · Yaoyao Zhong, Junbin Xiao, Wei Ji, Yicong Li 외

Video Question Answering (VideoQA) aims to answer natural language questions according to the given videos. It has earned increasing attention with recent research trends in joint vision and language understanding. Yet, …

Question AnsweringVideo Question Answering

Modality-Inconsistent Continual Learning of Multimodal Large Language Models

2024-12-17 · Weiguo Pian, Shijian Deng, Shentong Mo, Yunhui Guo 외

In this paper, we introduce Modality-Inconsistent Continual Learning (MICL), a new continual learning scenario for Multimodal Large Language Models (MLLMs) that involves tasks with inconsistent modalities (image, audio, …

Continual LearningKnowledge DistillationQuestion Answering

AcademicGPT: Empowering Academic Research

2023-11-21 · Shufa Wei, Xiaolong Xu, Xianbiao Qi, Xi Yin 외

Large Language Models (LLMs) have demonstrated exceptional capabilities across various natural language processing tasks. Yet, many of these advanced LLMs are tailored for broad, general-purpose applications. In this tec…

Abstract generationGeneral KnowledgeMMLUQuestion Answering

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

2025-01-21 · Yi Wang, Xinhao Li, Ziang Yan, Yinan He 외

This paper aims to improve the performance of video multimodal large language models (MLLM) via long and rich context (LRC) modeling. As a result, we develop a new version of InternVideo2.5 with a focus on enhancing the …

Object TrackingReferring Expression SegmentationReferring Video Object SegmentationVideo Understanding