paper-with-me

홈 › Papers

Data-Efficient Multimodal Fusion on a Single GPU

2023-12-15 · CVPR 2024 1 · Noël Vouitsis, Zhaoyan Liu, Satya Krishna Gorti, Valentin Villecroze, Jesse C. Cresswell, Guangwei Yu, Gabriel Loaiza-Ganem, Maksims Volkovs

The goal of multimodal alignment is to learn a single latent space that is shared between multimodal inputs. The most powerful models in this space have been trained using massive datasets of paired inputs and large-scale computational resources, making them prohibitively expensive to train in many practical scenarios. We surmise that existing unimodal encoders pre-trained on large amounts of unimodal data should provide an effective bootstrap to create multimodal models from unimodal ones at much lower costs. We therefore propose FuseMix, a multimodal augmentation scheme that operates on the latent spaces of arbitrary pre-trained unimodal encoders. Using FuseMix for multimodal alignment, we achieve competitive performance -- and in certain cases outperform state-of-the art methods -- in both image-text and audio-text retrieval, with orders of magnitude less compute and data: for example, we outperform CLIP on the Flickr30K text-to-image retrieval task with $\sim \! 600\times$ fewer GPU days and $\sim \! 80\times$ fewer image-text pairs. Additionally, we show how our method can be applied to convert pre-trained text-to-image generative models into audio-to-image ones. Code is available at: https://github.com/layer6ai-labs/fusemix.

📄 PDF Abstract BibTeX arXiv:2312.10144

Code (2)

layer6ai-labs/fusemix 공식 구현 pytorch
layer6ai-labs/direct-cms pytorch

Tasks

GPUImage RetrievalRetrievalText Retrieval

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Defending Multimodal Fusion Models against Single-Source Adversaries

2022-06-25 · CVPR 2021 1 · Karren Yang, Wan-Yi Lin, Manash Barman, Filipe Condessa 외

Beyond achieving high performance across many vision tasks, multimodal models are expected to be robust to single-source faults due to the availability of redundant information between modalities. In this paper, we inves…

Action Recognitionobject-detectionObject DetectionSentiment Analysis

Learning Deep Multimodal Feature Representation with Asymmetric Multi-layer Fusion

2021-08-11 · Yikai Wang, Fuchun Sun, Ming Lu, Anbang Yao

We propose a compact and effective framework to fuse multimodal features at multiple layers in a single network. The framework consists of two innovative fusion schemes. Firstly, unlike existing multimodal methods that n…

Representation LearningSemantic SegmentationTranslation

A review of deep learning-based information fusion techniques for multimodal medical image classification

2024-04-23 · Yihao Li, Mostafa El Habib Daho, Pierre-Henri Conze, Rachid Zeghlache 외

Multimodal medical imaging plays a pivotal role in clinical diagnosis and research, as it combines information from various imaging modalities to provide a more comprehensive understanding of the underlying pathology. Re…

image-classificationImage ClassificationManagementMedical Image Classification

MMDR: A Result Feature Fusion Object Detection Approach for Autonomous System

2023-04-19 · Wendong Zhang

Object detection has been extensively utilized in autonomous systems in recent years, encompassing both 2D and 3D object detection. Recent research in this field has primarily centered around multimodal approaches for ad…

3D Object DetectionObjectobject-detectionObject Detection

Bone-conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models

2026-01-18 · Sina Khanagha, Bunlong Lay, Timo Gerkmann arxiv

Single-channel speech enhancement models face significant performance degradation in extremely noisy environments. While prior work has shown that complementary bone-conducted speech can guide enhancement, effective inte…

Speech Enhancement