paper-with-me

Papers

Learning to Embed Multi-Modal Contexts for Situated Conversational Agents

2022-01-16 · ACL ARR January 2022 1 · Anonymous

The Situated Interactive Multi-Modal Conversations (SIMMC) 2.0 aims to create virtual shopping assistants that can accept complex multi-modal inputs, i.e. visual appearances of objects and user utterances. It consists of four subtasks, multi-modal disambiguation (MM-Disamb), multi-modal coreference resolution (MM-Coref), multi-modal dialog state tracking (MM-DST), and response retrieval and generation. While many task-oriented dialog systems usually tackle each subtask separately, we propose a jointly learned multi-modal encoder-decoder that incorporates visual inputs and performs all four subtasks at once for efficiency. This approach won the MM-Coref and response retrieval subtasks and nominated runner-up for the remaining subtasks using a single unified model at the 10th Dialog Systems Technology Challenge (DSTC10), setting a high bar for the novel task of multi-modal task-oriented dialog systems.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

coreference-resolutionCoreference ResolutionDecoderdialog state trackingResponse GenerationRetrieval

Similar Papers 제목 키워드 기반

Learning to Embed Multi-Modal Contexts for Situated Conversational Agents

2022-07-01 · Findings (NAACL) 2022 7 · Haeju Lee, Oh Joon Kwon, Yunseon Choi, Minho Park 외

The Situated Interactive Multi-Modal Conversations (SIMMC) 2.0 aims to create virtual shopping assistants that can accept complex multi-modal inputs, i.e. visual appearances of objects and user utterances. It consists of…

coreference-resolutionCoreference ResolutionDecoderdialog state tracking+3

Which One Are You Referring To? Multimodal Object Identification in Situated Dialogue

2023-02-28 · Holy Lovenia, Samuel Cahyawijaya, Pascale Fung

The demand for multimodal dialogue systems has been rising in various domains, emphasizing the importance of interpreting multimodal inputs from conversational and situational contexts. We explore three methods to tackle…

SIMMC: Situated Interactive Multi-Modal Conversational Data Collection And Evaluation Platform

2019-11-07 · Paul A. Crook, Shivani Poddar, Ankita De, Semir Shafi 외

As digital virtual assistants become ubiquitous, it becomes increasingly important to understand the situated behaviour of users as they interact with these assistants. To this end, we introduce SIMMC, an extension to Pa…

AI AgentUnity

Towards Objective Evaluation of Socially-Situated Conversational Robots: Assessing Human-Likeness through Multimodal User Behaviors

2023-08-21 · Koji Inoue, Divesh Lala, Keiko Ochi, Tatsuya Kawahara 외

This paper tackles the challenging task of evaluating socially situated conversational robots and presents a novel objective evaluation approach that relies on multimodal user behaviors. In this study, our main focus is …

Situated and Interactive Multimodal Conversations

2020-06-02 · COLING 2020 8 · Seungwhan Moon, Satwik Kottur, Paul A. Crook, Ankita De 외

Next generation virtual assistants are envisioned to handle multimodal inputs (e.g., vision, memories of previous interactions, in addition to the user's utterances), and perform multimodal actions (e.g., displaying a ro…

Response Generation