paper-with-me

Papers

InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback

2025-05-29 · Boyuan Chen, Donghai Hong, Jiaming Ji, Jiacheng Zheng, Bowen Dong, Jiayi Zhou, Kaile Wang, Juntao Dai, Xuyao Wang, Wenqi Chen, Qirui Zheng, Wenxin Li, Sirui Han, Yike Guo, Yaodong Yang

As multimodal large models (MLLMs) continue to advance across challenging tasks, a key question emerges: What essential capabilities are still missing? A critical aspect of human learning is continuous interaction with the environment -- not limited to language, but also involving multimodal understanding and generation. To move closer to human-level intelligence, models must similarly support multi-turn, multimodal interaction. In particular, they should comprehend interleaved multimodal contexts and respond coherently in ongoing exchanges. In this work, we present an initial exploration through the InterMT -- the first preference dataset for multi-turn multimodal interaction, grounded in real human feedback. In this exploration, we particularly emphasize the importance of human oversight, introducing expert annotations to guide the process, motivated by the fact that current MLLMs lack such complex interactive capabilities. InterMT captures human preferences at both global and local levels into nine sub-dimensions, consists of 15.6k prompts, 52.6k multi-turn dialogue instances, and 32.4k human-labeled preference pairs. To compensate for the lack of capability for multi-modal understanding and generation, we introduce an agentic workflow that leverages tool-augmented MLLMs to construct multi-turn QA instances. To further this goal, we introduce InterMT-Bench to assess the ability of MLLMs in assisting judges with multi-turn, multimodal tasks. We demonstrate the utility of \InterMT through applications such as judge moderation and further reveal the multi-turn scaling law of judge model. We hope the open-source of our data can help facilitate further research on aligning current MLLMs to the next step. Our project website can be found at https://pku-intermt.github.io .

📄 PDF Abstract BibTeX arXiv:2505.23950

Code (0)

등록된 구현이 없습니다.

Tasks

multimodal interaction

Similar Papers 제목 키워드 기반

Regularized Conditional Diffusion Model for Multi-Task Preference Alignment

2024-04-07 · Xudong Yu, Chenjia Bai, Haoran He, Changhong Wang 외

Sequential decision-making is desired to align with human intents and exhibit versatility across various tasks. Previous methods formulate it as a conditional generation process, utilizing return-conditioned diffusion mo…

D4RLDecision MakingSequential Decision Making

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation

2026-02-06 · Wenxun Dai, Zhiyuan Zhao, Yule Zhong, Yiji Cheng 외 arxiv

Unified multimodal models (UMMs) have achieved remarkable progress yet remain constrained by a single-turn interaction paradigm, effectively functioning as solvers for independent requests rather than assistants in conti…

Text-to-Image Generationmultimodal generation

Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image

2025-12-18 · Yushi Hu, Reyhane Askari-Hemmat, Melissa Hall, Emily Dinan 외 arxiv

Reward models (RMs) are essential for training large language models (LLMs), but remain underexplored for omni models that handle interleaved image and text sequences. We introduce Multimodal RewardBench 2 (MMRB2), the f…

Multimodal ReasoningImage Editing

Aligning LLMs with Individual Preferences via Interaction

2024-10-04 · Shujin Wu, May Fung, Cheng Qian, Jeonghwan Kim 외

As large language models (LLMs) demonstrate increasingly advanced capabilities, aligning their behaviors with human values and preferences becomes crucial for their wide adoption. While previous research focuses on gener…

TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving

2025-08-10 · Wenqian Cui, Lei Zhu, Xiaohui Li, Zhihan Guo 외 arxiv

Full-Duplex Speech Language Models (FD-SLMs) are specialized foundation models designed to enable natural, real-time spoken interactions by modeling complex conversational turn-taking such as interruptions, backchannels,…