paper-with-me

홈 › Papers

Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal Models

2023-01-16 · CVPR 2023 1 · Zhiqiu Lin, Samuel Yu, Zhiyi Kuang, Deepak Pathak, Deva Ramanan

The ability to quickly learn a new task with minimal instruction - known as few-shot learning - is a central aspect of intelligent agents. Classical few-shot benchmarks make use of few-shot samples from a single modality, but such samples may not be sufficient to characterize an entire concept class. In contrast, humans use cross-modal information to learn new concepts efficiently. In this work, we demonstrate that one can indeed build a better ${\bf visual}$ dog classifier by ${\bf read}$ing about dogs and ${\bf listen}$ing to them bark. To do so, we exploit the fact that recent multimodal foundation models such as CLIP learn cross-modal encoders that map different modalities to the same representation space. Specifically, we propose a simple strategy for ${\bf cross-modal}$ ${\bf adaptation}$: we treat examples from different modalities as additional few-shot examples. For example, by simply repurposing class names as an additional training sample, we trivially turn any n-shot learning problem into a (n+1)-shot problem. This allows us to produce SOTA results with embarrassingly simple linear classifiers. We show that our approach can be combined with existing methods such as prefix tuning, adapters, and classifier ensembling. Finally, to explore other modalities beyond vision and language, we construct the first (to our knowledge) audiovisual few-shot benchmark and use cross-modal training to improve the performance of both image and audio classification.

📄 PDF Abstract BibTeX arXiv:2301.06267

Code (1)

linzhiqiu/cross_modal_adaptation 공식 구현 pytorch

Tasks

Audio ClassificationFew-Shot Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Reliable Multimodality Eye Disease Screening via Mixture of Student's t Distributions

2023-03-17 · Ke Zou, Tian Lin, Xuedong Yuan, Haoyu Chen 외

Multimodality eye disease screening is crucial in ophthalmology as it integrates information from diverse sources to complement their respective performances. However, the existing methods are weak in assessing the relia…

Decision Making

TCT: A Cross-supervised Learning Method for Multimodal Sequence Representation

2019-10-23 · Wubo Li, Wei Zou, Xiangang Li

Multimodalities provide promising performance than unimodality in most tasks. However, learning the semantic of the representations from multimodalities efficiently is extremely challenging. To tackle this, we propose th…

Foundation Models in Remote Sensing: Evolving from Unimodality to Multimodality

2026-03-01 · Danfeng Hong, Chenyu Li, Xuyang Li, Gustau Camps-Valls 외 arxiv

Remote sensing (RS) techniques are increasingly crucial for deepening our understanding of the planet. As the volume and diversity of RS data continue to grow exponentially, there is an urgent need for advanced data mode…

Rethinking Multimodal Content Moderation from an Asymmetric Angle with Mixed-modality

2023-05-17 · Jialin Yuan, Ye Yu, Gaurav Mittal, Matthew Hall 외

There is a rapidly growing need for multimodal content moderation (CM) as more and more content on social media is multimodal in nature. Existing unimodal CM systems may fail to catch harmful content that crosses modalit…

Modality for Scenario Analysis and Maximum Likelihood Allocation

2020-05-06 · Takaaki Koike, Marius Hofert

We study the variability of a risk from the statistical viewpoint of multimodality of the conditional loss distribution given that the aggregate loss equals an exogenously provided capital. This conditional distribution …