Generative Visual Dialogue System via Adaptive Reasoning and Weighted Likelihood Estimation
The key challenge of generative Visual Dialogue (VD) systems is to respond to human queries with informative answers in natural and contiguous conversation flow. Traditional Maximum Likelihood Estimation (MLE)-based methods only learn from positive responses but ignore the negative responses, and consequently tend to yield safe or generic responses. To address this issue, we propose a novel training scheme in conjunction with weighted likelihood estimation (WLE) method. Furthermore, an adaptive multi-modal reasoning module is designed, to accommodate various dialogue scenarios automatically and select relevant information accordingly. The experimental results on the VisDial benchmark demonstrate the superiority of our proposed algorithm over other state-of-the-art approaches, with an improvement of 5.81% on recall@10.
Code (0)
등록된 구현이 없습니다.
Tasks
Visual DialogSimilar Papers 제목 키워드 기반
Scene-Aware Conversational ADAS with Generative AI for Real-Time Driver Assistance
While autonomous driving technologies continue to advance, current Advanced Driver Assistance Systems (ADAS) remain limited in their ability to interpret scene context or engage with drivers through natural language. The…
Autonomous DrivingKBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual Dialogue
Visual dialogue is a challenging task that needs to extract implicit information from both visual (image) and textual (dialogue history) contexts. Classical approaches pay more attention to the integration of the current…
Information RetrievalRetrievalEMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
Speech emotions play a crucial role in human-computer interaction, shaping engagement and context-aware communication. Despite recent advances in spoken dialogue systems, a holistic system for evaluating emotional reason…
DVD: A Diagnostic Dataset for Multi-step Reasoning in Video Grounded Dialogue
A video-grounded dialogue system is required to understand both dialogue, which contains semantic dependencies from turn to turn, and video, which contains visual cues of spatial and temporal scene variations. Building s…
DiagnosticObject TrackingVisual ReasoningAre You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning
The Visual Dialogue task requires an agent to engage in a conversation about an image with a human. It represents an extension of the Visual Question Answering task in that the agent needs to answer a question about an i…
Question AnsweringReinforcement LearningVisual DialogVisual Question Answering+1