paper-with-me

Papers

Revisit Multimodal Meta-Learning through the Lens of Multi-Task Learning

2021-10-27 · NeurIPS 2021 12 · Milad Abdollahzadeh, Touba Malekzadeh, Ngai-Man Cheung

Multimodal meta-learning is a recent problem that extends conventional few-shot meta-learning by generalizing its setup to diverse multimodal task distributions. This setup makes a step towards mimicking how humans make use of a diverse set of prior skills to learn new skills. Previous work has achieved encouraging performance. In particular, in spite of the diversity of the multimodal tasks, previous work claims that a single meta-learner trained on a multimodal distribution can sometimes outperform multiple specialized meta-learners trained on individual unimodal distributions. The improvement is attributed to knowledge transfer between different modes of task distributions. However, there is no deep investigation to verify and understand the knowledge transfer between multimodal tasks. Our work makes two contributions to multimodal meta-learning. First, we propose a method to quantify knowledge transfer between tasks of different modes at a micro-level. Our quantitative, task-level analysis is inspired by the recent transference idea from multi-task learning. Second, inspired by hard parameter sharing in multi-task learning and a new interpretation of related work, we propose a new multimodal meta-learner that outperforms existing work by considerable margins. While the major focus is on multimodal meta-learning, our work also attempts to shed light on task interaction in conventional meta-learning. The code for this project is available at https://miladabd.github.io/KML.

📄 PDF Abstract BibTeX arXiv:2110.14202

Code (1)

sutd-visual-computing-group/KML-Classification 공식 구현 pytorch

Tasks

Meta-LearningMulti-Task LearningTransfer Learning

Similar Papers 제목 키워드 기반

Hellinger Multimodal Variational Autoencoders

2026-01-10 · Huyen Vo, Isabel Valera arxiv

Multimodal variational autoencoders (VAEs) are widely used for weakly supervised generative learning with multiple modalities. Predominant methods aggregate unimodal inference distributions using either a product of expe…

Learned split-spectrum metalens for obstruction-free broadband imaging in the visible

2026-01-27 · Seungwoo Yoon, Dohyun Kang, Eunsue Choi, Sohyun Lee 외 arxiv

Obstructions such as raindrops, fences, or dust degrade captured images, especially when mechanical cleaning is infeasible. Conventional solutions to obstructions rely on a bulky compound optics array or computational in…

Semantic SegmentationObject Detection

Revisiting MLLM Token Technology through the Lens of Classical Visual Coding

2025-08-19 · Jinming Liu, Junyan Lin, Yuntao Wei, Kele Shao 외 arxiv

Classical visual coding and Multimodal Large Language Model (MLLM) token technology share the core objective - maximizing information fidelity while minimizing computational cost. Therefore, this paper reexamines MLLM to…

Segmenting the Complex and Irregular in Two-Phase Flows: A Real-World Empirical Study with SAM2

2025-08-07 · Semanur Küçük, Cosimo Della Santina, Angeliki Laskari arxiv

Segmenting gas bubbles in multiphase flows is a critical yet unsolved challenge in numerous industrial settings, from metallurgical processing to maritime drag reduction. Traditional approaches-and most recent learning-b…

Transfer Learning

Feature Fusion Revisited: Multimodal CTR Prediction for MMCTR Challenge

2025-04-26 · Junjie Zhou

With the rapid advancement of Multimodal Large Language Models (MLLMs), an increasing number of researchers are exploring their application in recommendation systems. However, the high latency associated with large model…

Click-Through Rate PredictionInformation RetrievalPredictionRecommendation Systems+2