Response to "Visual Dialogue without Vision or Dialogue" (Massiceti et al., 2018)
In a recent workshop paper, Massiceti et al. presented a baseline model and subsequent critique of Visual Dialog (Das et al., CVPR 2017) that raises what we believe to be unfounded concerns about the dataset and evaluation. This article intends to rebut the critique and clarify potential confusions for practitioners and future participants in the Visual Dialog challenge.
Code (0)
등록된 구현이 없습니다.
Tasks
Visual DialogSimilar Papers 제목 키워드 기반
A Visually-grounded First-person Dialogue Dataset with Verbal and Non-verbal Responses
In real-world dialogue, first-person visual information about where the other speakers are and what they are paying attention to is crucial to understand their intentions. Non-verbal responses also play an important role…
Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation
In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an ava…
ChatbotLanguage ModelingLanguage ModellingLarge Language ModelVisualizing Dialogues: Enhancing Image Selection through Dialogue Understanding with Large Language Models
Recent advancements in dialogue systems have highlighted the significance of integrating multimodal responses, which enable conveying ideas through diverse modalities rather than solely relying on text-based interactions…
Dialogue UnderstandingImage RetrievalRetrievalEx-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coord…
Text is NOT Enough: Integrating Visual Impressions into Open-domain Dialogue Generation
Open-domain dialogue generation in natural language processing (NLP) is by default a pure-language task, which aims to satisfy human need for daily communication on open-ended topics by producing related and informative …
DecoderDialogue GenerationDialogue Understanding