paper-with-me

홈 › Papers

Text is NOT Enough: Integrating Visual Impressions into Open-domain Dialogue Generation

2021-09-13 · Lei Shen, Haolan Zhan, Xin Shen, Yonghao Song, Xiaofang Zhao

Open-domain dialogue generation in natural language processing (NLP) is by default a pure-language task, which aims to satisfy human need for daily communication on open-ended topics by producing related and informative responses. In this paper, we point out that hidden images, named as visual impressions (VIs), can be explored from the text-only data to enhance dialogue understanding and help generate better responses. Besides, the semantic dependency between an dialogue post and its response is complicated, e.g., few word alignments and some topic transitions. Therefore, the visual impressions of them are not shared, and it is more reasonable to integrate the response visual impressions (RVIs) into the decoder, rather than the post visual impressions (PVIs). However, both the response and its RVIs are not given directly in the test process. To handle the above issues, we propose a framework to explicitly construct VIs based on pure-language dialogue datasets and utilize them for better dialogue understanding and generation. Specifically, we obtain a group of images (PVIs) for each post based on a pre-trained word-image mapping model. These PVIs are used in a co-attention encoder to get a post representation with both visual and textual information. Since the RVIs are not provided directly during testing, we design a cascade decoder that consists of two sub-decoders. The first sub-decoder predicts the content words in response, and applies the word-image mapping model to get those RVIs. Then, the second sub-decoder generates the response based on the post and RVIs. Experimental results on two open-domain dialogue datasets show that our proposed approach achieves superior performance over competitive baselines.

📄 PDF Abstract BibTeX arXiv:2109.05778

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDialogue GenerationDialogue Understanding

Similar Papers 제목 키워드 기반

Impressions: Understanding Visual Semiotics and Aesthetic Impact

2023-10-27 · Julia Kruk, Caleb Ziems, Diyi Yang

Is aesthetic impact different from beauty? Is visual salience a reflection of its capacity for effective communication? We present Impressions, a novel dataset through which to investigate the semiotics of images, and ho…

Image CaptioningImage Description

Characterizing Hirability via Personality and Behavior

2020-06-22 · Harshit Malik, Hersh Dhillon, Roland Goecke, Ramanathan Subramanian

While personality traits have been extensively modeled as behavioral constructs, we model \textbf{\textit{job hirability}} as a \emph{personality construct}. On the {\emph{First Impressions Candidate Screening}} (FICS) d…

Evaluation of GPT-4 for chest X-ray impression generation: A reader study on performance and perception

2023-11-12 · Sebastian Ziegelmayer, Alexander W. Marka, Nicolas Lenhart, Nadja Nehls 외

The remarkable generative capabilities of multimodal foundation models are currently being explored for a variety of applications. Generating radiological impressions is a challenging task that could significantly reduce…

Impressions2Font: Generating Fonts by Specifying Impressions

2021-03-18 · Seiya Matsuda, Akisato Kimura, Seiichi Uchida

Various fonts give us various impressions, which are often represented by words. This paper proposes Impressions2Font (Imp2Font) that generates font images with specific impressions. Imp2Font is an extended version of co…

Which Parts Determine the Impression of the Font?

2021-03-26 · Masaya Ueda, Akisato Kimura, Seiichi Uchida

Various fonts give different impressions, such as legible, rough, and comic-text.This paper aims to analyze the correlation between the local shapes, or parts, and the impression of fonts. By focusing on local shapes ins…

regression