Scene-Aware Dialogue
1개 벤치마크 · 논문 8편 · 이 태스크의 논문 보기 →
Benchmarks
AVSD
Most implemented
Audio-Visual Scene-Aware Dialog
An Embodied Generalist Agent in 3D World
Maintaining Common Ground in Dynamic Environments
A Simple Baseline for Audio-Visual Scene-Aware Dialog
Papers
An Embodied Generalist Agent in 3D World
Leveraging massive knowledge from large language models (LLMs), recent machine learning models show notable successes in general-purpose task solving in diverse domains such as computer vision and robotics. However, seve…
3D dense captioning3D Question Answering (3D-QA)Question AnsweringRobot Manipulation+3Maintaining Common Ground in Dynamic Environments
Common grounding is the process of creating and maintaining mutual understandings, which is a critical aspect of sophisticated human communication. While various task settings have been proposed in existing literature, t…
End-To-End Dialogue ModellingGoal-Oriented Dialogue SystemsScene-Aware DialogueMultimodal Dialogue State Tracking By QA Approach with Data Augmentation
Recently, a more challenging state tracking task, Audio-Video Scene-Aware Dialogue (AVSD), is catching an increasing amount of attention among researchers. Different from purely text-based dialogue state tracking, the di…
Data AugmentationDecoderDialogue State TrackingOpen-Domain Question Answering+2Multi-step Joint-Modality Attention Network for Scene-Aware Dialogue System
Understanding dynamic scenes and dialogue contexts in order to converse with users has been challenging for multimodal dialogue systems. The 8-th Dialog System Technology Challenge (DSTC8) proposed an Audio Visual Scene-…
Scene-Aware DialogueEntropy-Enhanced Multimodal Attention Model for Scene-Aware Dialogue Generation
With increasing information from social media, there are more and more videos available. Therefore, the ability to reason on a video is important and deserves to be discussed. TheDialog System Technology Challenge (DSTC7…
Dialogue GenerationScene-Aware DialogueReactive Multi-Stage Feature Fusion for Multimodal Dialogue Modeling
Visual question answering and visual dialogue tasks have been increasingly studied in the multimodal field towards more practical real-world scenarios. A more challenging task, audio visual scene-aware dialogue (AVSD), i…
Question AnsweringScene-Aware DialogueVisual DialogVisual Question Answering+1