paper-with-me

홈 › Papers

Relating by Contrasting: A Data-efficient Framework for Multimodal Generative Models

2020-07-02 · ICLR 2021 1 · Yuge Shi, Brooks Paige, Philip H. S. Torr, N. Siddharth

Multimodal learning for generative models often refers to the learning of abstract concepts from the commonality of information in multiple modalities, such as vision and language. While it has proven effective for learning generalisable representations, the training of such models often requires a large amount of "related" multimodal data that shares commonality, which can be expensive to come by. To mitigate this, we develop a novel contrastive framework for generative model learning, allowing us to train the model not just by the commonality between modalities, but by the distinction between "related" and "unrelated" multimodal data. We show in experiments that our method enables data-efficient multimodal learning on challenging datasets for various multimodal VAE models. We also show that under our proposed framework, the generative model can accurately identify related samples from unrelated ones, making it possible to make use of the plentiful unlabeled, unpaired multimodal data.

📄 PDF Abstract BibTeX arXiv:2007.01179

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

USD Coin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Multimodal Physiological Signals Representation Learning via Multiscale Contrasting for Depression Recognition

2024-06-22 · Kai Shao, Rui Wang, Yixue Hao, Long Hu 외

Depression recognition based on physiological signals such as functional near-infrared spectroscopy (fNIRS) and electroencephalogram (EEG) has made considerable progress. However, most existing studies ignore the complem…

Data AugmentationEEGElectroencephalogram (EEG)Representation Learning+2

Multimodal Scientific Learning Beyond Diffusions and Flows

2026-02-01 · Leonardo Ferreira Guilhoto, Akshat Kaushal, Paris Perdikaris arxiv

Scientific machine learning (SciML) increasingly requires models that capture multimodal conditional uncertainty arising from ill-posed inverse problems, multistability, and chaotic dynamics. While recent work has favore…

Decoding Predictive Inference in Visual Language Processing via Spatiotemporal Neural Coherence

2025-12-24 · Sean C. Borneman, Julia Krebs, Ronnie B. Wilbur, Evie A. Malaia arxiv

Human language processing relies on the brain's capacity for predictive inference. We present a machine learning framework for decoding neural (EEG) responses to dynamic visual language stimuli in Deaf signers. Using coh…

A Hybrid Deterministic Framework for Named Entity Extraction in Broadcast News Video

2026-02-09 · Andrea Filiberto Lucas, Dylan Seychell arxiv

The growing volume of video-based news content has heightened the need for transparent and reliable methods to extract on-screen information. Yet the variability of graphical layouts, typographic conventions, and platfor…

Information Extraction

Investigating Inner Properties of Multimodal Representation and Semantic Compositionality with Brain-based Componential Semantics

2017-11-15 · Shaonan Wang, Jiajun Zhang, Nan Lin, Cheng-qing Zong

Multimodal models have been proven to outperform text-based approaches on learning semantic representations. However, it still remains unclear what properties are encoded in multimodal representations, in what aspects do…

Learning Semantic RepresentationsNatural Language Understanding