Tackling Situated Multi-Modal Task-Oriented Dialogs with a Single Transformer Model
The Situated Interactive Multi-Modal Conversations (SIMMC) 2.0 aims to create virtual shopping assistants that can accept complex multi-modal inputs, i.e. visual appearances of objects and user utterances. It consists of four subtasks, multi-modal disambiguation (MM-Disamb), multi-modal coreference resolution (MM-Coref), multi-modal dialog state tracking (MM-DST), and response retrieval and generation. While many task-oriented dialog systems usually tackle each subtask separately, we propose a jointly learned encoder-decoder that performs all four subtasks at once for efficiency. Moreover, we handle the multi-modality of the challenge by representing visual objects as special tokens whose joint embedding is learned via auxiliary tasks. This approach won the MM-Coref and response retrieval subtasks and nominated runner-up for the remaining subtasks using a single unified model. In particular, our model achieved 81.5\% MRR, 71.2\% R@1, 95.0\% R@5, 98.2\% R@10, and 1.9 mean rank in response retrieval task, setting a high bar for the state-of-the-art result in the SIMMC 2.0 track of the Dialog Systems Technology Challenge 10 (DSTC10).
Code (0)
등록된 구현이 없습니다.
Tasks
coreference-resolutionCoreference ResolutionDecoderdialog state trackingRetrievalSimilar Papers 제목 키워드 기반
Learning to Embed Multi-Modal Contexts for Situated Conversational Agents
The Situated Interactive Multi-Modal Conversations (SIMMC) 2.0 aims to create virtual shopping assistants that can accept complex multi-modal inputs, i.e. visual appearances of objects and user utterances. It consists of…
coreference-resolutionCoreference ResolutionDecoderdialog state tracking+2Learning to Embed Multi-Modal Contexts for Situated Conversational Agents
The Situated Interactive Multi-Modal Conversations (SIMMC) 2.0 aims to create virtual shopping assistants that can accept complex multi-modal inputs, i.e. visual appearances of objects and user utterances. It consists of…
coreference-resolutionCoreference ResolutionDecoderdialog state tracking+3SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations
Next generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal environment. Existing task-oriented dialog…
DiversityLanguage ModelingLanguage ModellingOverview of the Ninth Dialog System Technology Challenge: DSTC9
This paper introduces the Ninth Dialog System Technology Challenge (DSTC-9). This edition of the DSTC focuses on applying end-to-end dialog technologies for four distinct tasks in dialog systems, namely, 1. Task-oriented…
Interactive Evaluation of DialogMulti-modal Situated Reasoning in 3D Scenes
Situation awareness is essential for understanding and reasoning about 3D scenes in embodied AI agents. However, existing datasets and benchmarks for situated understanding are limited in data modality, diversity, scale,…
3D Question Answering (3D-QA)