paper-with-me

홈 › Papers

CL-CrossVQA: A Continual Learning Benchmark for Cross-Domain Visual Question Answering

2022-11-19 · Yao Zhang, Haokun Chen, Ahmed Frikha, Yezi Yang, Denis Krompass, Gengyuan Zhang, Jindong Gu, Volker Tresp

Visual Question Answering (VQA) is a multi-discipline research task. To produce the right answer, it requires an understanding of the visual content of images, the natural language questions, as well as commonsense reasoning over the information contained in the image and world knowledge. Recently, large-scale Vision-and-Language Pre-trained Models (VLPMs) have been the mainstream approach to VQA tasks due to their superior performance. The standard practice is to fine-tune large-scale VLPMs pre-trained on huge general-domain datasets using the domain-specific VQA datasets. However, in reality, the application domain can change over time, necessitating VLPMs to continually learn and adapt to new domains without forgetting previously acquired knowledge. Most existing continual learning (CL) research concentrates on unimodal tasks, whereas a more practical application scenario, i.e, CL on cross-domain VQA, has not been studied. Motivated by this, we introduce CL-CrossVQA, a rigorous Continual Learning benchmark for Cross-domain Visual Question Answering, through which we conduct extensive experiments on 4 VLPMs, 4 CL approaches, and 5 VQA datasets from different domains. In addition, by probing the forgetting phenomenon of the intermediate layers, we provide insights into how model architecture affects CL performance, why CL approaches can help mitigate forgetting in VLPMs to some extent, and how to design CL approaches suitable for VLPMs in this challenging continual learning environment. To facilitate future work on CL for cross-domain VQA, we will release our datasets and code.

📄 PDF Abstract BibTeX arXiv:2211.10567

Code (0)

등록된 구현이 없습니다.

Tasks

Continual LearningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)World Knowledge

Similar Papers 제목 키워드 기반

CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA Generalization

2021-11-01 · EMNLP 2021 11 · Arjun Akula, Soravit Changpinyo, Boqing Gong, Piyush Sharma 외

One challenge in evaluating visual question answering (VQA) models in the cross-dataset adaptation setting is that the distribution shifts are multi-modal, making it difficult to identify if it is the shifts in visual or…

Answer GenerationQuestion-Answer-GenerationQuestion AnsweringVisual Question Answering+1

Audio-Visual Continual Test-Time Adaptation without Forgetting

2026-02-20 · Sarthak Kumar Maharana, Akshay Mehra, Bhavya Ramakrishna, Yunhui Guo 외 arxiv

Audio-visual continual test-time adaptation involves continually adapting a source audio-visual model at test-time, to unlabeled non-stationary domains, where either or both modalities can be distributionally shifted, wh…

Test-time Adaptation

Can You Hear, Localize, and Segment Continually? An Exemplar-Free Continual Learning Benchmark for Audio-Visual Segmentation

2026-03-09 · Siddeshwar Raghavan, Gautham Vinod, Bruce Coburn, Fengqing Zhu arxiv

Audio-Visual Segmentation (AVS) aims to produce pixel-level masks of sound producing objects in videos, by jointly learning from audio and visual signals. However, real-world environments are inherently dynamic, causing …

Continual Learning

Decorate the Newcomers: Visual Domain Prompt for Continual Test Time Adaptation

2022-12-08 · Yulu Gan, Yan Bai, Yihang Lou, Xianzheng Ma 외

Continual Test-Time Adaptation (CTTA) aims to adapt the source model to continually changing unlabeled target domains without access to the source data. Existing methods mainly focus on model-based adaptation in a self-t…

Prompt LearningTest-time Adaptation

Class-Incremental Grouping Network for Continual Audio-Visual Learning

2023-09-11 · ICCV 2023 1 · Shentong Mo, Weiguo Pian, Yapeng Tian

Continual learning is a challenging problem in which models need to be trained on non-stationary data across sequential tasks for class-incremental learning. While previous methods have focused on using either regulariza…

audio-visual learningclass-incremental learningClass Incremental LearningContinual Learning+3