paper-with-me

Papers

Multimodal Transformer with Pointer Network for the DSTC8 AVSD Challenge

2020-02-25 · Hung Le, Nancy F. Chen

Audio-Visual Scene-Aware Dialog (AVSD) is an extension from Video Question Answering (QA) whereby the dialogue agent is required to generate natural language responses to address user queries and carry on conversations. This is a challenging task as it consists of video features of multiple modalities, including text, visual, and audio features. The agent also needs to learn semantic dependencies among user utterances and system responses to make coherent conversations with humans. In this work, we describe our submission to the AVSD track of the 8th Dialogue System Technology Challenge. We adopt dot-product attention to combine text and non-text features of input video. We further enhance the generation capability of the dialogue agent by adopting pointer networks to point to tokens from multiple source sequences in each generation step. Our systems achieve high performance in automatic metrics and obtain 5th and 6th place in human evaluation among all submissions.

📄 PDF Abstract BibTeX arXiv:2002.10695

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVideo Question Answering

Methods 이 논문이 사용한 방법론

Six Ways To Communicate To Someone At Expedia Via Phone And Email's. To communicate or get human at Expedia, the quickest option is typically to call their customer service at +1-888-829-0881 or +1(805) 330 (4056). You can also use the live chat…

Similar Papers 제목 키워드 기반

Audio-Visual Scene-Aware Dialog and Reasoning using Audio-Visual Transformers with Joint Student-Teacher Learning

2021-10-13 · Ankit P. Shah, Shijie Geng, Peng Gao, Anoop Cherian 외

In previous work, we have proposed the Audio-Visual Scene-Aware Dialog (AVSD) task, collected an AVSD dataset, developed AVSD technologies, and hosted an AVSD challenge track at both the 7th and 8th Dialog System Technol…

Region Proposal

Bridging Text and Video: A Universal Multimodal Transformer for Video-Audio Scene-Aware Dialog

2020-02-01 · Zekang Li, Zongjia Li, Jinchao Zhang, Yang Feng 외

Audio-Visual Scene-Aware Dialog (AVSD) is a task to generate responses when chatting about a given video, which is organized as a track of the 8th Dialog System Technology Challenge (DSTC8). To solve the task, we propose…

Dialogue GenerationMulti-Task LearningText Generation

DSTC8-AVSD: Multimodal Semantic Transformer Network with Retrieval Style Word Generator

2020-04-01 · Hwanhee Lee, Seunghyun Yoon, Franck Dernoncourt, Doo Soon Kim 외

Audio Visual Scene-aware Dialog (AVSD) is the task of generating a response for a question with a given scene, video, audio, and the history of previous turns in the dialog. Existing systems for this task employ the tran…

DecoderRetrievalWord Embeddings

Audio Visual Scene-Aware Dialog Generation with Transformer-based Video Representations

2022-02-21 · Yoshihiro Yamazaki, Shota Orihashi, Ryo Masumura, Mihiro Uchida 외

There have been many attempts to build multimodal dialog systems that can respond to a question about given audio-visual information, and the representative task for such systems is the Audio Visual Scene-Aware Dialog (A…

Answer GenerationVideo Understanding

Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7

2018-06-01 · Huda Alamri, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das 외

Scene-aware dialog systems will be able to have conversations with users about the objects and events around them. Progress on such systems can be made by integrating state-of-the-art technologies from multiple research …

Video DescriptionVisual Dialog