paper-with-me

홈 › Papers

X-Fusion: Introducing New Modality to Frozen Large Language Models

2025-04-29 · Sicheng Mo, Thao Nguyen, Xun Huang, Siddharth Srinivasan Iyer, Yijun Li, Yuchen Liu, Abhishek Tandon, Eli Shechtman, Krishna Kumar Singh, Yong Jae Lee, Bolei Zhou, Yuheng Li

We propose X-Fusion, a framework that extends pretrained Large Language Models (LLMs) for multimodal tasks while preserving their language capabilities. X-Fusion employs a dual-tower design with modality-specific weights, keeping the LLM's parameters frozen while integrating vision-specific information for both understanding and generation. Our experiments demonstrate that X-Fusion consistently outperforms alternative architectures on both image-to-text and text-to-image tasks. We find that incorporating understanding-focused data improves generation quality, reducing image data noise enhances overall performance, and feature alignment accelerates convergence for smaller models but has minimal impact on larger ones. Our findings provide valuable insights into building efficient unified multimodal models.

📄 PDF Abstract BibTeX arXiv:2504.20996

Code (0)

등록된 구현이 없습니다.

Tasks

Image to text

Similar Papers 제목 키워드 기반

Diffusion-Link: Diffusion Probabilistic Model for Bridging the Audio-Text Modality Gap

2025-10-13 · KiHyun Nam, Jongmin Choi, Hyeongkeun Lee, Jungwoo Heo 외 arxiv

Contrastive audio-language pretraining yields powerful joint representations, yet a persistent audio-text modality gap limits the benefits of coupling multimodal encoders with large language models (LLMs). We present Dif…

Audio captioning

Parameter-Efficient Modality-Balanced Symmetric Fusion for Multimodal Remote Sensing Semantic Segmentation

2026-03-18 · Haocheng Li, Juepeng Zheng, Shuangxi Miao, Ruibo Lu 외 arxiv

Multimodal remote sensing semantic segmentation enhances scene interpretation by exploiting complementary physical cues from heterogeneous data. Although pretrained Vision Foundation Models (VFMs) provide strong general-…

Semantic Segmentation

Cross-lingual Matryoshka Representation Learning across Speech and Text

2026-02-23 · Yaya Sy, Dioula Doucouré, Christophe Cerisara, Irina Illina arxiv

Speakers of under-represented languages face both a language barrier, as most online knowledge is in a few dominant languages, and a modality barrier, since information is largely text-based while many languages are prim…

Representation LearningIntent Detection

Robust Cross-Modal Foundation Model Perception for Underwater Robots under Degraded Visual Conditions

2026-08-20 · Mohammad Arif Ul Alam arxiv

Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation, low illumination, scattering, and blur. Although sonar provides complementary i…

Unsupervised Learning for Missing Modalities in Multimodal Learning

2026-06-14 · Hassan Ismkhan, Hamid Bouchahcia arxiv

This paper addresses the missing-modality challenge in multi-modal learning by introducing Unsupervised Learning for Missing Modalities in Multi-Modal Learning (UL4M4), a flexible framework that imputes missing feature e…