paper-with-me

Papers

Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs

2025-10-10 · Yumin Choi, Dongki Kim, Jinheon Baek, Sung Ju Hwang arxiv

Large Language Models (LLMs) have shown remarkable success, and their multimodal expansions (MLLMs) further unlock capabilities spanning images, videos, and other modalities beyond text. However, despite this shift, prompt optimization approaches, designed to reduce the burden of manual prompt crafting while maximizing performance, remain confined to text, ultimately limiting the full potential of MLLMs. Motivated by this gap, we introduce the new problem of multimodal prompt optimization, which expands the prior definition of prompt optimization to the multimodal space defined by the pairs of textual and non-textual prompts. To tackle this problem, we then propose the Multimodal Prompt Optimizer (MPO), a unified framework that not only performs the joint optimization of multimodal prompts through alignment-preserving updates but also guides the selection process of candidate prompts by leveraging earlier evaluations as priors in a Bayesian-based selection strategy. Through extensive experiments across diverse modalities that go beyond text, such as images, videos, and even molecules, we demonstrate that MPO outperforms leading text-only optimization methods, establishing multimodal prompt optimization as a crucial step to realizing the potential of MLLMs.

📄 PDF Abstract BibTeX arXiv:2510.09201

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

YoChameleon: Personalized Vision and Language Generation

2025-04-29 · Thao Nguyen, Krishna Kumar Singh, Jing Shi, Trung Bui 외

Large Multimodal Models (e.g., GPT-4, Gemini, Chameleon) have evolved into powerful tools with millions of users. However, they remain generic models and lack personalized knowledge of specific user concepts. Previous wo…

Image GenerationText Generation

POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models

2024-06-06 · Jianben He, Xingbo Wang, Shiyi Liu, Guande Wu 외

Large language models (LLMs) have exhibited impressive abilities for multimodal content comprehension and reasoning with proper prompting in zero- or few-shot settings. Despite the proliferation of interactive systems de…

Multimodal ReasoningPrompt Engineering

Yo'Chameleon: Personalized Vision and Language Generation

2025-01-01 · CVPR 2025 1 · Thao Nguyen, Krishna Kumar Singh, Jing Shi, Trung Bui 외

Large Multimodal Models (e.g., GPT-4, Gemini, Chameleon) have evolved into powerful tools with millions of users. However, they remain generic models and lack personalized knowledge of specific user concepts. Previou…

Image GenerationText Generation

Decoupled Multimodal Prototypes for Visual Recognition with Missing Modalities

2025-05-13 · Jueqing Lu, Yuanyuan Qi, Xiaohao Yang, Shujie Zhou 외

Multimodal learning enhances deep learning models by enabling them to perceive and understand information from multiple data modalities, such as visual and textual inputs. However, most existing approaches assume the ava…

Distilled Prompt Learning for Incomplete Multimodal Survival Prediction

2025-01-01 · CVPR 2025 1 · Yingxue Xu, Fengtao Zhou, Chenyu Zhao, Yihui Wang 외

The integration of multimodal data including pathology images and gene profiles is widely applied in precise survival prediction. Despite recent advances in multimodal survival models, collecting complete modalities …

PredictionPrompt LearningSurvival Prediction