paper-with-me

Papers

Audio Visual Scene-Aware Dialog Generation with Transformer-based Video Representations

2022-02-21 · Yoshihiro Yamazaki, Shota Orihashi, Ryo Masumura, Mihiro Uchida, Akihiko Takashima

There have been many attempts to build multimodal dialog systems that can respond to a question about given audio-visual information, and the representative task for such systems is the Audio Visual Scene-Aware Dialog (AVSD). Most conventional AVSD models adopt the Convolutional Neural Network (CNN)-based video feature extractor to understand visual information. While a CNN tends to obtain both temporally and spatially local information, global information is also crucial for boosting video understanding because AVSD requires long-term temporal visual dependency and whole visual information. In this study, we apply the Transformer-based video feature that can capture both temporally and spatially global representations more efficiently than the CNN-based feature. Our AVSD model with its Transformer-based feature attains higher objective performance scores for answer generation. In addition, our model achieves a subjective score close to that of human answers in DSTC10. We observed that the Transformer-based visual feature is beneficial for the AVSD task because our model tends to correctly answer the questions that need a temporally and spatially broad range of visual information.

📄 PDF Abstract BibTeX arXiv:2202.09979

Code (0)

등록된 구현이 없습니다.

Tasks

Answer GenerationVideo Understanding

Similar Papers 제목 키워드 기반

Bridging Text and Video: A Universal Multimodal Transformer for Video-Audio Scene-Aware Dialog

2020-02-01 · Zekang Li, Zongjia Li, Jinchao Zhang, Yang Feng 외

Audio-Visual Scene-Aware Dialog (AVSD) is a task to generate responses when chatting about a given video, which is organized as a track of the 8th Dialog System Technology Challenge (DSTC8). To solve the task, we propose…

Dialogue GenerationMulti-Task LearningText Generation

Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7

2018-06-01 · Huda Alamri, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das 외

Scene-aware dialog systems will be able to have conversations with users about the objects and events around them. Progress on such systems can be made by integrating state-of-the-art technologies from multiple research …

Video DescriptionVisual Dialog

Audio-Visual Scene-Aware Dialog

2019-01-25 · Huda Alamri, Vincent Cartillier, Abhishek Das, Jue Wang 외

We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To ans…

Scene-Aware Dialogue

Audio Visual Scene-Aware Dialog

2019-06-01 · CVPR 2019 6 · Huda Alamri, Vincent Cartillier, Abhishek Das, Jue Wang 외

We introduce the task of scene-aware dialog. Our goal is to generate a complete and natural response to a question about a scene, given video and audio of the scene and the history of previous turns in the dialog. To ans…

Leveraging Topics and Audio Features with Multimodal Attention for Audio Visual Scene-Aware Dialog

2019-12-20 · Shachi H. Kumar, Eda Okur, Saurav Sahay, Jonathan Huang 외

With the recent advancements in Artificial Intelligence (AI), Intelligent Virtual Assistants (IVA) such as Alexa, Google Home, etc., have become a ubiquitous part of many homes. Currently, such IVAs are mostly audio-base…

Audio ClassificationResponse Generation