paper-with-me

홈 › Papers

Response to "Visual Dialogue without Vision or Dialogue" (Massiceti et al., 2018)

2019-01-16 · Abhishek Das, Devi Parikh, Dhruv Batra

In a recent workshop paper, Massiceti et al. presented a baseline model and subsequent critique of Visual Dialog (Das et al., CVPR 2017) that raises what we believe to be unfounded concerns about the dataset and evaluation. This article intends to rebut the critique and clarify potential confusions for practitioners and future participants in the Visual Dialog challenge.

📄 PDF Abstract BibTeX arXiv:1901.05531

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Dialog

Similar Papers 제목 키워드 기반

A Visually-grounded First-person Dialogue Dataset with Verbal and Non-verbal Responses

2020-11-01 · EMNLP 2020 11 · Hisashi Kamezawa, Noriki Nishida, Nobuyuki Shimizu, Takashi Miyazaki 외

In real-world dialogue, first-person visual information about where the other speakers are and what they are paying attention to is crucial to understand their intentions. Non-verbal responses also play an important role…

Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

2024-06-12 · Se Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim 외

In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an ava…

ChatbotLanguage ModelingLanguage ModellingLarge Language Model

Visualizing Dialogues: Enhancing Image Selection through Dialogue Understanding with Large Language Models

2024-07-04 · Chang-Sheng Kao, Yun-Nung Chen

Recent advancements in dialogue systems have highlighted the significance of integrating multimodal responses, which enable conveying ideas through diverse modalities rather than solely relying on text-based interactions…

Dialogue UnderstandingImage RetrievalRetrieval

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

2026-08-11 · Haoyu Zhang, Zhipeng Li, Xiaoying Tang, Tianshu Yu 외 hf

Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coord…

Text is NOT Enough: Integrating Visual Impressions into Open-domain Dialogue Generation

2021-09-13 · Lei Shen, Haolan Zhan, Xin Shen, Yonghao Song 외

Open-domain dialogue generation in natural language processing (NLP) is by default a pure-language task, which aims to satisfy human need for daily communication on open-ended topics by producing related and informative …

DecoderDialogue GenerationDialogue Understanding