paper-with-me

Papers

Efficient Multimodal Fusion via Interactive Prompting

2023-04-13 · CVPR 2023 1 · Yaowei Li, Ruijie Quan, Linchao Zhu, Yi Yang

Large-scale pre-training has brought unimodal fields such as computer vision and natural language processing to a new era. Following this trend, the size of multi-modal learning models constantly increases, leading to an urgent need to reduce the massive computational cost of finetuning these models for downstream tasks. In this paper, we propose an efficient and flexible multimodal fusion method, namely PMF, tailored for fusing unimodally pre-trained transformers. Specifically, we first present a modular multimodal fusion framework that exhibits high flexibility and facilitates mutual interactions among different modalities. In addition, we disentangle vanilla prompts into three types in order to learn different optimizing objectives for multimodal learning. It is also worth noting that we propose to add prompt vectors only on the deep layers of the unimodal transformers, thus significantly reducing the training memory usage. Experiment results show that our proposed method achieves comparable performance to several other multimodal finetuning methods with less than 3% trainable parameters and up to 66% saving of training memory usage.

📄 PDF Abstract BibTeX arXiv:2304.06306

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

2024-06-28 · Yuxuan Zhang, Tianheng Cheng, Rui Hu, Lei Liu 외

Segment Anything Model (SAM) has attracted widespread attention for its superior interactive segmentation capabilities with visual prompts while lacking further exploration of text prompts. In this paper, we empirically …

Interactive SegmentationLanguage ModelingLanguage ModellingReferring Expression+2

Chat2Layout: Interactive 3D Furniture Layout with a Multimodal LLM

2024-07-31 · Can Wang, Hongliang Zhong, Menglei Chai, Mingming He 외

Automatic furniture layout is long desired for convenient interior design. Leveraging the remarkable visual reasoning capabilities of multimodal large language models (MLLMs), recent methods address layout generation in …

In-Context LearningLayout DesignLayout GenerationVisual Prompting+1

SAMSA 2.0: Prompting Segment Anything with Spectral Angles for Hyperspectral Interactive Medical Image Segmentation

2025-08-01 · Alfie Roddan, Tobias Czempiel, Chi Xu, Daniel S. Elson 외 arxiv

We present SAMSA 2.0, an interactive segmentation framework for hyperspectral medical imaging that introduces spectral angle prompting to guide the Segment Anything Model (SAM) using spectral similarity alongside spatial…

Medical Image SegmentationInteractive Segmentation

Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

2023-10-17 · Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou 외

We present Set-of-Mark (SoM), a new visual prompting method, to unleash the visual grounding abilities of large multimodal models (LMMs), such as GPT-4V. As illustrated in Fig. 1 (right), we employ off-the-shelf interact…

Interactive SegmentationReferring ExpressionReferring Expression ComprehensionSegmentation+2

TSAM: Temporal SAM Augmented with Multimodal Prompts for Referring Audio-Visual Segmentation

2025-01-01 · CVPR 2025 1 · Abduljalil Radman, Jorma Laaksonen

Referring audio-visual segmentation (Ref-AVS) aims to segment objects within audio-visual scenes using multimodal cues embedded in text expressions. While the Segment Anything Model (SAM) has revolutionized visual se…

Referring Audio-Visual Segmentation