paper-with-me

홈 › Papers

On the Computational Benefit of Multimodal Learning

2023-09-25 · Zhou Lu

Human perception inherently operates in a multimodal manner. Similarly, as machines interpret the empirical world, their learning processes ought to be multimodal. The recent, remarkable successes in empirical multimodal learning underscore the significance of understanding this paradigm. Yet, a solid theoretical foundation for multimodal learning has eluded the field for some time. While a recent study by Lu (2023) has shown the superior sample complexity of multimodal learning compared to its unimodal counterpart, another basic question remains: does multimodal learning also offer computational advantages over unimodal learning? This work initiates a study on the computational benefit of multimodal learning. We demonstrate that, under certain conditions, multimodal learning can outpace unimodal learning exponentially in terms of computation. Specifically, we present a learning task that is NP-hard for unimodal learning but is solvable in polynomial time by a multimodal algorithm. Our construction is based on a novel modification to the intersection of two half-spaces problem.

📄 PDF Abstract BibTeX arXiv:2309.13782

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Online Convolutional Dictionary Learning for Multimodal Imaging

2017-06-13 · Kevin Degraux, Ulugbek S. Kamilov, Petros T. Boufounos, Dehong Liu

Computational imaging methods that can exploit multiple modalities have the potential to enhance the capabilities of traditional sensing systems. In this paper, we propose a new method that reconstructs multimodal images…

Dictionary Learning

Sparse Fusion for Multimodal Transformers

2021-11-23 · Yi Ding, Alex Rich, Mason Wang, Noah Stier 외

Multimodal classification is a core task in human-centric machine learning. We observe that information is highly complementary across modalities, thus unimodal information can be drastically sparsified prior to multimod…

Re-evaluating the Need for Multimodal Signals in Unsupervised Grammar Induction

2022-12-20 · Boyi Li, Rodolfo Corona, Karttikeya Mangalam, Catherine Chen 외

Are multimodal inputs necessary for grammar induction? Recent work has shown that multimodal training inputs can improve grammar induction. However, these improvements are based on comparisons to weak text-only baselines…

Constituency Parsing

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation

2025-07-10 · Yupu Liang, Yaping Zhang, Zhiyang Zhang, Yang Zhao 외 arxiv

Document Image Machine Translation (DIMT) aims to translate text within document images, facing generalization challenges due to limited training data and the complex interplay between visual and textual information. To …

Computational EfficiencyDomain GeneralizationMachine Translation

Synergy vs. Noise: Performance-Guided Multimodal Fusion For Biochemical Recurrence-Free Survival in Prostate Cancer

2025-11-14 · Seth Alain Chang, Muhammad Mueez Amjad, Noorul Wahab, Ethar Alzaid 외 arxiv

Multimodal deep learning (MDL) has emerged as a transformative approach in computational pathology. By integrating complementary information from multiple data sources, MDL models have demonstrated superior predictive pe…

Multimodal Deep Learning