paper-with-me

홈 › Papers

Associative Conversation Model: Generating Visual Information from Textual Information

2018-01-01 · ICLR 2018 1 · Yoichi Ishibashi, Hisashi Miyamori

In this paper, we propose the Associative Conversation Model that generates visual information from textual information and uses it for generating sentences in order to utilize visual information in a dialogue system without image input. In research on Neural Machine Translation, there are studies that generate translated sentences using both images and sentences, and these studies show that visual information improves translation performance. However, it is not possible to use sentence generation algorithms using images for the dialogue systems since many text-based dialogue systems only accept text input. Our approach generates (associates) visual information from input text and generates response text using context vector fusing associative visual information and sentence textual information. A comparative experiment between our proposed model and a model without association showed that our proposed model is generating useful sentences by associating visual information related to sentences. Furthermore, analysis experiment of visual association showed that our proposed model generates (associates) visual information effective for sentence generation.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

Associative Multichannel Autoencoder for Multimodal Word Representation

2018-10-01 · EMNLP 2018 10 · Shaonan Wang, Jiajun Zhang, Cheng-qing Zong

In this paper we address the problem of learning multimodal word representations by integrating textual, visual and auditory inputs. Inspired by the re-constructive and associative nature of human memory, we propose a no…

Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech

2025-09-18 · Taesoo Kim, Yongsik Jo, Hyunmin Song, Taehwan Kim arxiv

Human conversation involves language, speech, and visual cues, with each medium providing complementary information. For instance, speech conveys a vibe or tone not fully captured by text alone. While multimodal LLMs foc…

MultiDM-GCN: Aspect-guided Response Generation in Multi-domain Multi-modal Dialogue System using Graph Convolutional Network

2020-11-01 · Findings of the Association for Computational Linguistics 2020 · Mauajama Firdaus, Nidhi Thakur, Asif Ekbal

In the recent past, dialogue systems have gained immense popularity and have become ubiquitous. During conversations, humans not only rely on languages but seek contextual information through visual contents as well. In …

DecoderResponse Generation

What Should I Ask? Using Conversationally Informative Rewards for Goal-Oriented Visual Dialog

2019-07-28 · Pushkar Shukla, Carlos Elmadjian, Richika Sharan, Vivek Kulkarni 외

The ability to engage in goal-oriented conversations has allowed humans to gain knowledge, reduce uncertainty, and perform tasks more efficiently. Artificial agents, however, are still far behind humans in having goal-dr…

Reinforcement LearningVisual Dialog

What Should I Ask? Using Conversationally Informative Rewards for Goal-oriented Visual Dialog.

2019-07-01 · ACL 2019 7 · Pushkar Shukla, Carlos Elmadjian, Richika Sharan, Vivek Kulkarni 외

The ability to engage in goal-oriented conversations has allowed humans to gain knowledge, reduce uncertainty, and perform tasks more efficiently. Artificial agents, however, are still far behind humans in having goal-dr…

Reinforcement LearningVisual Dialog