paper-with-me

Papers

Modular and Parameter-Efficient Multimodal Fusion with Prompting

2022-03-15 · Findings (ACL) 2022 5 · Sheng Liang, Mengjie Zhao, Hinrich Schütze

Recent research has made impressive progress in large-scale multimodal pre-training. In the context of the rapid growth of model size, it is necessary to seek efficient and flexible methods other than finetuning. In this paper, we propose to use prompt vectors to align the modalities. Our method achieves comparable performance to several other multimodal fusion methods in low-resource settings. We further show that our method is modular and parameter-efficient for processing tasks involving two or more data modalities.

📄 PDF Abstract BibTeX arXiv:2203.08055

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Modular and Parameter-Efficient Multimodal Fusion with Prompting

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Recent research has made impressive progress in large-scale multimodal pre-training. In the context of the rapid growth of model size, it is necessary to seek efficient and flexible methods other than fine-tuning. In thi…

Conditional Prompt Tuning for Multimodal Fusion

2023-11-28 · Ruixiang Jiang, Lingbo Liu, Changwen Chen

We show that the representation of one modality can effectively guide the prompting of another modality for parameter-efficient multimodal fusion. Specifically, we first encode one modality and use its representation as …

Multimodal Modular Chain of Thoughts in Energy Performance Certificate Assessment

2026-02-20 · Zhen Peng, Peter J. Bentley arxiv

Accurate evaluation of building energy performance remains challenging in regions where scalable Energy Performance Certificate (EPC) assessments are unavailable. This paper presents a cost-efficient framework that lever…

Efficient Multimodal Fusion via Interactive Prompting

2023-04-13 · CVPR 2023 1 · Yaowei Li, Ruijie Quan, Linchao Zhu, Yi Yang

Large-scale pre-training has brought unimodal fields such as computer vision and natural language processing to a new era. Following this trend, the size of multi-modal learning models constantly increases, leading to an…

ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way

2025-07-11 · Rajarshi Roy, Devleena Das, Ankesh Banerjee, Arjya Bhattacharjee 외

We introduce ByDeWay, a training-free framework designed to enhance the performance of Multimodal Large Language Models (MLLMs). ByDeWay uses a novel prompting strategy called Layered-Depth-Based Prompting (LDP), which i…

Depth EstimationHallucinationLanguage ModelingLanguage Modelling+2