paper-with-me

Papers

Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives

2025-11-23 · Kai Jiang, Siqi Huang, Xiangyu Chen, Jiawei Shao, Hongyuan Zhang, Ping Luo, Xuelong Li arxiv

Multimodal large language models (MLLMs) deployed on devices must adapt to continuously changing visual scenarios such as variations in background and perspective, to effectively perform complex visual tasks. To investigate catastrophic forgetting under real-world scenario shifts, we construct a multimodal visual understanding dataset (MSVQA), covering four distinct scenarios and perspectives: high-altitude, underwater, low-altitude, and indoor environments. Furthermore, we propose UNIFIER (mUltimodal coNtInual learning with MLLMs From multi-scenarIo pERspectives), a continual learning (CL) framework designed to address visual discrepancies while learning different scenarios. Compared to existing CL methods, UNIFIER enables knowledge accumulation within the same scenario and mutual enhancement across different scenarios via Vision Representation Expansion (VRE) and Vision Consistency Constraint (VCC). Experimental results show that UNIFIER improves the last-step VQA scores by 2.70%~10.62% and the last-step F1 scores by 3.40%~7.69% compared to the state-of-the-art method, QUAD, in 20-step cross-scenario continual learning tasks. MSVQA dataset is available at https://huggingface.co/datasets/Kaij00/MSVQA.

📄 PDF Abstract BibTeX arXiv:2511.18507

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Learning

Similar Papers 제목 키워드 기반

Continual-NExT: A Unified Comprehension And Generation Continual Learning Framework

2026-02-20 · Jingyang Qiao, Zhizhong Zhang, Xin Tan, Jingyu Gong 외 arxiv

Dual-to-Dual MLLMs refer to Multimodal Large Language Models, which can enable unified multimodal comprehension and generation through text and image modalities. Although exhibiting strong instantaneous learning and gene…

Continual Learning

A Practitioner's Guide to Continual Multimodal Pretraining

2024-08-26 · Karsten Roth, Vishaal Udandarao, Sebastian Dziadzio, Ameya Prabhu 외

Multimodal foundation models serve numerous applications at the intersection of vision and language. Still, despite being pretrained on extensive data, they become outdated over time. To keep models updated, research int…

Continual LearningContinual PretrainingMeta-Learning

Safety of Multimodal Large Language Models on Images and Texts

2024-02-01 · Xin Liu, Yichen Zhu, Yunshi Lan, Chao Yang 외

Attracted by the impressive power of Multimodal Large Language Models (MLLMs), the public is increasingly utilizing them to improve the efficiency of daily work. Nonetheless, the vulnerabilities of MLLMs to unsafe instru…

Survey

MIKE: A New Benchmark for Fine-grained Multimodal Entity Knowledge Editing

2024-02-18 · Jiaqi Li, Miaozeng Du, Chuanyi Zhang, Yongrui Chen 외

Multimodal knowledge editing represents a critical advancement in enhancing the capabilities of Multimodal Large Language Models (MLLMs). Despite its potential, current benchmarks predominantly focus on coarse-grained kn…

knowledge editing

MLLM-CL: Continual Learning for Multimodal Large Language Models

2025-06-05 · Hongbo Zhao, Fei Zhu, Rundong Wang, Gaofeng Meng 외

Recent Multimodal Large Language Models (MLLMs) excel in vision-language understanding but face challenges in adapting to dynamic real-world scenarios that require continuous integration of new knowledge and skills. Whil…

Continual Learning