paper-with-me

홈 › Papers

Are You Talking to Me? Reasoned Visual Dialog Generation through Adversarial Learning

2017-11-21 · CVPR 2018 6 · Qi Wu, Peng Wang, Chunhua Shen, Ian Reid, Anton Van Den Hengel

The Visual Dialogue task requires an agent to engage in a conversation about an image with a human. It represents an extension of the Visual Question Answering task in that the agent needs to answer a question about an image, but it needs to do so in light of the previous dialogue that has taken place. The key challenge in Visual Dialogue is thus maintaining a consistent, and natural dialogue while continuing to answer questions correctly. We present a novel approach that combines Reinforcement Learning and Generative Adversarial Networks (GANs) to generate more human-like responses to questions. The GAN helps overcome the relative paucity of training data, and the tendency of the typical MLE-based approach to generate overly terse answers. Critically, the GAN is tightly integrated into the attention mechanism that generates human-interpretable reasons for each answer. This means that the discriminative model of the GAN has the task of assessing whether a candidate answer is generated by a human or not, given the provided reason. This is significant because it drives the generative model to produce high quality answers that are well supported by the associated reasoning. The method also generates the state-of-the-art results on the primary benchmark.

📄 PDF Abstract BibTeX arXiv:1711.07613

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringReinforcement LearningVisual DialogVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dogecoin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Conversational Human Audio-visual Talking Dialogue Generation

2026-07-02 · Junhao Song, Lluis Guasch, Xilin He, Zhongyu Yang 외 arxiv

Large-scale dyadic interactive audio-visual dialogue (DIAD) datasets provide fundamental data resources for developing humanoid interactive virtual agents and digital humans. However, collecting such data is time-consumi…

Dialogue Generation

When to Talk: Chatbot Controls the Timing of Talking during Multi-turn Open-domain Dialogue Generation

2019-12-20 · Tian Lan, Xian-Ling Mao, He-Yan Huang, Wei Wei

Despite the multi-turn open-domain dialogue systems have attracted more and more attention and made great progress, the existing dialogue systems are still very boring. Nearly all the existing dialogue models only provid…

ChatbotDialogue Generation

TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation

2025-12-23 · Ji-Hoon Kim, Junseok Ahn, Doyeop Kwak, Joon Son Chung 외 arxiv

The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conversational systems, recent studies have …

Dialogue Generation

Personalized Large Vision-Language Models

2024-12-23 · Chau Pham, Hoang Phan, David Doermann, Yunjie Tian

The personalization model has gained significant attention in image generation yet remains underexplored for large vision-language models (LVLMs). Beyond generic ones, with personalization, LVLMs handle interactive dialo…

Image Generation

Talking Face Generation by Adversarially Disentangled Audio-Visual Representation

2018-07-20 · Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo 외

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the s…

Face GenerationLip ReadingRetrievalTalking Face Generation+1