paper-with-me

홈 › Papers

Sample-efficient Integration of New Modalities into Large Language Models

2025-09-04 · Osman Batur İnce, André F. T. Martins, Oisin Mac Aodha, Edoardo M. Ponti arxiv

Multimodal foundation models can process several modalities. However, since the space of possible modalities is large and evolving over time, training a model from scratch to encompass all modalities is unfeasible. Moreover, integrating a modality into a pre-existing foundation model currently requires a significant amount of paired data, which is often not available for low-resource modalities. In this paper, we introduce a method for sample-efficient modality integration (SEMI) into Large Language Models (LLMs). To this end, we devise a hypernetwork that can adapt a shared projector -- placed between modality-specific encoders and an LLM -- to any modality. The hypernetwork, trained on high-resource modalities (i.e., text, speech, audio, video), is conditioned on a few samples from any arbitrary modality at inference time to generate a suitable adapter. To increase the diversity of training modalities, we artificially multiply the number of encoders through isometric transformations. We find that SEMI achieves a significant boost in sample efficiency during few-shot integration of new modalities (i.e., satellite images, astronomical images, inertial measurements, and molecules) with encoders of arbitrary embedding dimensionality. For instance, to reach the same accuracy as 32-shot SEMI, training the projector from scratch needs 64$\times$ more data. As a result, SEMI holds promise to extend the modality coverage of foundation models.

📄 PDF Abstract BibTeX arXiv:2509.04606

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

2024-02-19 · Jun Zhan, Junqi Dai, Jiasheng Ye, Yunhua Zhou 외

We introduce AnyGPT, an any-to-any multimodal language model that utilizes discrete representations for the unified processing of various modalities, including speech, text, images, and music. AnyGPT can be trained stabl…

Language ModelingLanguage ModellingLarge Language Model

When Large Language Models Meet Speech: A Survey on Integration Approaches

2025-02-26 · Zhengdong Yang, Shuichiro Shimizu, Yahan Yu, Chenhui Chu

Recent advancements in large language models (LLMs) have spurred interest in expanding their application beyond text-based tasks. A large number of studies have explored integrating other modalities with LLMs, notably sp…

Efficiently Integrate Large Language Models with Visual Perception: A Survey from the Training Paradigm Perspective

2025-02-03 · Xiaorui Ma, Haoran Xie, S. Joe Qin

The integration of vision-language modalities has been a significant focus in multimodal learning, traditionally relying on Vision-Language Pretrained Models. However, with the advent of Large Language Models (LLMs), the…

Robust Multi-Omics Integration from Incomplete Modalities Significantly Improves Prediction of Alzheimer's Disease

2025-09-25 · Sungjoon Park, Kyungwook Lee, Soorin Yim, Doyeong Hwang 외 arxiv

Multi-omics data capture complex biomolecular interactions and provide insights into metabolism and disease. However, missing modalities hinder integrative analysis across heterogeneous omics. To address this, we present…

Feature Importance

ATFLRec: A Multimodal Recommender System with Audio-Text Fusion and Low-Rank Adaptation via Instruction-Tuned Large Language Model

2024-09-13 · Zezheng Qin

Recommender Systems (RS) play a pivotal role in boosting user satisfaction by providing personalized product suggestions in domains such as e-commerce and entertainment. This study examines the integration of multimodal …

Graph Neural NetworkLanguage ModelingLanguage ModellingLarge Language Model+2