paper-with-me

Papers

Multimodal Knowledge Expansion

2021-03-26 · ICCV 2021 10 · Zihui Xue, Sucheng Ren, Zhengqi Gao, Hang Zhao

The popularity of multimodal sensors and the accessibility of the Internet have brought us a massive amount of unlabeled multimodal data. Since existing datasets and well-trained models are primarily unimodal, the modality gap between a unimodal network and unlabeled multimodal data poses an interesting problem: how to transfer a pre-trained unimodal network to perform the same task on unlabeled multimodal data? In this work, we propose multimodal knowledge expansion (MKE), a knowledge distillation-based framework to effectively utilize multimodal data without requiring labels. Opposite to traditional knowledge distillation, where the student is designed to be lightweight and inferior to the teacher, we observe that a multimodal student model consistently denoises pseudo labels and generalizes better than its teacher. Extensive experiments on four tasks and different modalities verify this finding. Furthermore, we connect the mechanism of MKE to semi-supervised learning and offer both empirical and theoretical explanations to understand the denoising capability of a multimodal student.

📄 PDF Abstract BibTeX arXiv:2103.14431

Code (1)

zihuixue/mke 공식 구현

Tasks

DenoisingKnowledge DistillationSemantic Segmentation

Similar Papers 제목 키워드 기반

From Swath to Full-Disc: Advancing Precipitation Retrieval with Multimodal Knowledge Expansion

2025-06-08 · Zheng Wang, Kai Ying, Bin Xu, Chunjiao Wang 외

Accurate near-real-time precipitation retrieval has been enhanced by satellite-based technologies. However, infrared-based algorithms have low accuracy due to weak relations with surface precipitation, whereas passive mi…

Data IntegrationRetrieval

NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs

2026-08-08 · Jiayue Jin, Jingwei Zhang, Chen Wang, Jing Liu 외 hf

Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the language intelligence acquired during pretraining. In this work, we investigate this phenomenon from the …

PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents

2024-06-20 · Junjie Wang, Yin Zhang, Yatai Ji, Yuxiang Zhang 외

Recent advancements in Large Multimodal Models (LMMs) have leveraged extensive multimodal datasets to enhance capabilities in complex knowledge-driven tasks. However, persistent challenges in perceptual and reasoning err…

Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?

2026-04-03 · Qianshan Wei, Yishan Yang, Siyi Wang, Jinglin Chen 외 arxiv

Multimodal Large Language Models (MLLMs) are evolving from passive observers into active agents, solving problems through Visual Expansion (invoking visual tools) and Knowledge Expansion (open-web search). However, exist…

CoRE: Concept-Reasoning Expansion for Continual Brain Lesion Segmentation

2026-04-28 · Qianqian Chen, Anglin Liu, Jingyang Zhang, Yudong Zhang arxiv

Accurate brain lesion segmentation in MRI is vital for effective clinical diagnosis and treatment planning. Due to high annotation costs and strict data privacy regulations, universal models require employing Continual L…

Lesion SegmentationContinual Learning