paper-with-me

홈 › Papers

Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance

2025-09-06 · Weijie Shen, Xinrui Wang, Yuanqi Nie, Apiradee Boonmee arxiv

Current Large Language Models (LLMs) and Vision-Language Large Models (LVLMs) excel in single-turn tasks but face significant challenges in multi-turn interactions requiring deep contextual understanding and complex visual reasoning, often leading to fragmented reasoning, context loss, and hallucinations. To address these limitations, we propose Context-Aware Multi-Turn Visual Reasoning (CAMVR), a novel framework designed to empower LVLMs with robust and coherent multi-turn visual-textual inference capabilities. CAMVR introduces two key innovations: a Visual-Textual Context Memory Unit (VCMU), a dynamic read-write memory network that stores and manages critical visual features, textual semantic representations, and their cross-modal correspondences from each interaction turn; and an Adaptive Visual Focus Guidance (AVFG) mechanism, which leverages the VCMU's context to dynamically adjust the visual encoder's attention to contextually relevant image regions. Our multi-level reasoning integration strategy ensures that response generation is deeply coherent with both current inputs and accumulated historical context. Extensive experiments on challenging datasets, including VisDial, an adapted A-OKVQA, and our novel Multi-Turn Instruction Following (MTIF) dataset, demonstrate that CAMVR consistently achieves state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2509.05669

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingResponse GenerationVisual Reasoning

Similar Papers 제목 키워드 기반

MTMCS-Bench: Evaluating Contextual Safety of Multimodal Large Language Models in Multi-Turn Dialogues

2026-01-11 · Zheyuan Liu, Dongwhi Kim, Yixin Wan, Xiangchi Yuan 외 arxiv

Multimodal large language models (MLLMs) are increasingly deployed as assistants that interact through text and images, making it crucial to evaluate contextual safety when risk depends on both the visual scene and the e…

Intent Recognition

MMCR: Advancing Visual Language Model in Multimodal Multi-Turn Contextual Reasoning

2025-03-24 · Dawei Yan, Yang Li, Qing-Guo Chen, Weihua Luo 외

Compared to single-turn dialogue, multi-turn dialogue involving multiple images better aligns with the needs of real-world human-AI interactions. Additionally, as training data, it provides richer contextual reasoning in…

DiagnosticLanguage ModelingLanguage ModellingPrompt Engineering

OpenViDial: A Large-Scale, Open-Domain Dialogue Dataset with Visual Contexts

2020-12-30 · Yuxian Meng, Shuhe Wang, Qinghong Han, Xiaofei Sun 외

When humans converse, what a speaker will say next significantly depends on what he sees. Unfortunately, existing dialogue models generate dialogue utterances only based on preceding textual contexts, and visual contexts…

DecoderDialogue Generation

Router-Suggest: Dynamic Routing for Multimodal Auto-Completion in Visually-Grounded Dialogs

2026-01-09 · Sandeep Mishra, Devichand Budagam, Anubhab Mandal, Bishal Santra 외 arxiv

Real-time multimodal auto-completion is essential for digital assistants, chatbots, design tools, and healthcare consultations, where user inputs rely on shared visual context. We introduce Multimodal Auto-Completion (MA…

MILAB at SemEval-2019 Task 3: Multi-View Turn-by-Turn Model for Context-Aware Sentiment Analysis

2019-06-01 · SEMEVAL 2019 6 · Yoonhyung Lee, Yanghoon Kim, Kyomin Jung

This paper describes our system for SemEval-2019 Task 3: EmoContext, which aims to predict the emotion of the third utterance considering two preceding utterances in a dialogue. To address this challenge of predicting th…

Sentiment Analysis