paper-with-me

Papers

Beyond Modality Fusion: Deep Ensembles for Multimodal Classification

2026-07-06 · Ilya Burenko, Dmitry Vetrov arxiv

In multimodal classification, late-fusion approaches classify concatenated modality-specific features extracted by unimodal neural networks. When modality imbalance is pronounced, various regularization techniques have been proposed to balance the learning process and overcome the inferior performance of late-fusion networks. In contrast, this work demonstrates that multimodal data can be effectively classified without any explicit modality fusion, using deep ensembles of unimodal networks. We systematically compare deep ensembles to late-fusion networks at equal parameter count and show that ensembles consistently outperform state-of-the-art late-fusion methods designed to address modality imbalance. This advantage also holds over intermediate-fusion techniques we evaluated and over hybrid methods that combine unimodal and multimodal predictions. We propose and empirically validate a method for selecting the number of models per modality in an ensemble, avoiding computationally expensive exhaustive search. Under extreme modality imbalance and small ensemble sizes, the heuristic indicates that ensembles of unimodal models trained solely on the stronger modality are preferable; as the ensemble scales up, incorporating models from the weaker modality becomes beneficial. Both predictions align with our empirical findings. To systematically explore the challenges of optimizing multimodal models, we propose a synthetic multimodal framework that allows control over both the number of modalities and their predictive strength; our findings are consistent across synthetic and real-world datasets. Finally, by fitting scaling laws to bimodal datasets, we estimate the asymptotic performance of ensembles.

📄 PDF Abstract BibTeX arXiv:2607.05019

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap

2025-10-13 · KiHyun Nam, Jongmin Choi, Hyeongkeun Lee, Jungwoo Heo 외 arxiv

Contrastive audio-language pretraining yields powerful joint representations, yet a persistent audio-text modality gap limits the benefits of coupling multimodal encoders with large language models (LLMs). We present Dif…

Audio captioning

Multimodal Fusion Balancing Through Game-Theoretic Regularization

2024-11-11 · Konstantinos Kontras, Thomas Strypsteen, Christos Chatzichristos, Paul Pu Liang 외

Multimodal learning can complete the picture of information extraction by uncovering key dependencies between data sources. However, current systems fail to fully leverage multiple modalities for optimal performance. Thi…

Computational Efficiency

Multimodal Medical Image Classification via Synergistic Learning Pre-training

2025-09-22 · Qinghua Lin, Guang-Hai Liu, Zuoyong Li, Yang Li 외 arxiv

Multimodal pathological images are usually in clinical diagnosis, but computer vision-based multimodal image-assisted diagnosis faces challenges with modality fusion, especially in the absence of expert-annotated data. T…

Semi-supervised Medical Image ClassificationSelf-Supervised Learning

Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach

2026-07-16 · Elena Ryumina, Maxim Markitantov, Alexandr Axyonov, Fedor Shchetinin 외 arxiv

Automatic recognition of ambivalence and hesitancy is challenging because these states may be expressed through inconsistent linguistic, acoustic, facial, and contextual patterns, while top-performing systems often rely …

Multimodal Dynamics: Dynamical Fusion for Trustworthy Multimodal Classification

2022-01-01 · CVPR 2022 1 · Zongbo Han, Fan Yang, Junzhou Huang, Changqing Zhang 외

Integration of heterogeneous and high-dimensional data (e.g., multiomics) is becoming increasingly important. Existing multimodal classification algorithms mainly focus on improving performance by exploiting the comp…

ClassificationInformativenessMedical DiagnosisMulti-modal Classification