paper-with-me

Papers

Look Before you Speak: Visually Contextualized Utterances

2020-12-10 · CVPR 2021 1 · Paul Hongsuck Seo, Arsha Nagrani, Cordelia Schmid

While most conversational AI systems focus on textual dialogue only, conditioning utterances on visual context (when it's available) can lead to more realistic conversations. Unfortunately, a major challenge for incorporating visual context into conversational dialogue is the lack of large-scale labeled datasets. We provide a solution in the form of a new visually conditioned Future Utterance Prediction task. Our task involves predicting the next utterance in a video, using both visual frames and transcribed speech as context. By exploiting the large number of instructional videos online, we train a model to solve this task at scale, without the need for manual annotations. Leveraging recent advances in multimodal learning, our model consists of a novel co-attentional multimodal video transformer, and when trained on both textual and visual context, outperforms baselines that use textual inputs alone. Further, we demonstrate that our model trained for this task on unlabelled videos achieves state-of-the-art performance on a number of downstream VideoQA benchmarks such as MSRVTT-QA, MSVD-QA, ActivityNet-QA and How2QA.

📄 PDF Abstract BibTeX arXiv:2012.05710

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

COGMEN: COntextualized GNN based Multimodal Emotion recognitioN

2022-05-05 · NAACL 2022 7 · Abhinav Joshi, Ashwani Bhat, Ayush Jain, Atin Vikram Singh 외

Emotions are an inherent part of human interactions, and consequently, it is imperative to develop AI systems that understand and recognize human emotions. During a conversation involving various people, a person's emoti…

Emotion RecognitionEmotion Recognition in ConversationGraph Neural NetworkMultimodal Emotion Recognition

Filling the Gap of Utterance-aware and Speaker-aware Representation for Multi-turn Dialogue

2020-09-14 · Longxiang Liu, Zhuosheng Zhang, Hai Zhao, Xi Zhou 외

A multi-turn dialogue is composed of multiple utterances from two or more different speaker roles. Thus utterance- and speaker-aware clues are supposed to be well captured in models. However, in the existing retrieval-ba…

Retrieval

Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation

2026-08-23 · Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen arxiv

Conversational speech emotion recognition must reconcile acoustic evidence across temporal scales with two interaction processes: cross-speaker contextual influence and within-speaker emotion evolution. We propose DSSM-C…

Speech Emotion Recognition

Channel-aware Decoupling Network for Multi-turn Dialogue Comprehension

2023-01-10 · Zhuosheng Zhang, Hai Zhao, Longxiang Liu

Training machines to understand natural language and interact with humans is one of the major goals of artificial intelligence. Recent years have witnessed an evolution from matching networks to pre-trained language mode…

Speaking the Language of Your Listener: Audience-Aware Adaptation via Plug-and-Play Theory of Mind

2023-05-31 · Ece Takmaz, Nicolo' Brandizzi, Mario Giulianelli, Sandro Pezzelle 외

Dialogue participants may have varying levels of knowledge about the topic under discussion. In such cases, it is essential for speakers to adapt their utterances by taking their audience into account. Yet, it is an open…

Language ModelingLanguage ModellingOpen-Ended Question AnsweringText Generation