paper-with-me

Papers

Audio-Visual Scene-Aware Dialog and Reasoning using Audio-Visual Transformers with Joint Student-Teacher Learning

2021-10-13 · Ankit P. Shah, Shijie Geng, Peng Gao, Anoop Cherian, Takaaki Hori, Tim K. Marks, Jonathan Le Roux, Chiori Hori

In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at both the 7th and 8th Dialog System Technology Challenges (DSTC7, DSTC8). In these challenges, the best-performing systems relied heavily on human-generated descriptions of the video content, which were available in the datasets but would be unavailable in real-world applications. To promote further advancements for real-world applications, we proposed a third AVSD challenge, at DSTC10, with two modifications: 1) the human-created description is unavailable at inference time, and 2) systems must demonstrate temporal reasoning by finding evidence from the video to support each answer. This paper introduces the new task that includes temporal reasoning and our new extension of the AVSD dataset for DSTC10, for which we collected human-generated temporal reasoning data. We also introduce a baseline system built using an AV-transformer, which we released along with the new dataset. Finally, this paper introduces a new system that extends our baseline system with attentional multimodal fusion, joint student-teacher learning (JSTL), and model combination techniques, achieving state-of-the-art performances on the AVSD datasets for DSTC7, DSTC8, and DSTC10. We also propose two temporal reasoning methods for AVSD: one attention-based, and one based on a time-domain region proposal network.

📄 PDF Abstract BibTeX arXiv:2110.06894

Code (0)

등록된 구현이 없습니다.

Tasks

Region Proposal

Similar Papers 제목 키워드 기반

Multi-step Joint-Modality Attention Network for Scene-Aware Dialogue System

2020-01-17 · Yun-Wei Chu, Kuan-Yen Lin, Chao-Chun Hsu, Lun-Wei Ku

Understanding dynamic scenes and dialogue contexts in order to converse with users has been challenging for multimodal dialogue systems. The 8-th Dialog System Technology Challenge (DSTC8) proposed an Audio Visual Scene-…

Scene-Aware Dialogue

Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7

2018-06-01 · Huda Alamri, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das 외

Scene-aware dialog systems will be able to have conversations with users about the objects and events around them. Progress on such systems can be made by integrating state-of-the-art technologies from multiple research …

Video DescriptionVisual Dialog

Audio-Visual Scene-Aware Dialog

2019-01-25 · Huda Alamri, Vincent Cartillier, Abhishek Das, Jue Wang 외

We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To ans…

Scene-Aware Dialogue

Audio Visual Scene-Aware Dialog

2019-06-01 · CVPR 2019 6 · Huda Alamri, Vincent Cartillier, Abhishek Das, Jue Wang 외

We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To ans…

A Simple Baseline for Audio-Visual Scene-Aware Dialog

2019-04-11 · Idan Schwartz Alexander Schwing, Tamir Hazan and

The recently proposed audio-visual scene-aware dialog task paves the way to a more data-driven way of learning virtual assistants, smart speakers and car navigation systems. However, very little is known to date about ho…