paper-with-me

홈 › Papers

DBATES: DataBase of Audio features, Text, and visual Expressions in competitive debate Speeches

2021-03-26 · Taylan K. Sen, Gazi Naven, Luke Gerstner, Daryl Bagley, Raiyan Abdul Baten, Wasifur Rahman, Kamrul Hasan, Kurtis G. Haut, Abdullah Mamun, Samiha Samrose, Anne Solbu, R. Eric Barnes, Mark G. Frank, Ehsan Hoque

In this work, we present a database of multimodal communication features extracted from debate speeches in the 2019 North American Universities Debate Championships (NAUDC). Feature sets were extracted from the visual (facial expression, gaze, and head pose), audio (PRAAT), and textual (word sentiment and linguistic category) modalities of raw video recordings of competitive collegiate debaters (N=717 6-minute recordings from 140 unique debaters). Each speech has an associated competition debate score (range: 67-96) from expert judges as well as competitor demographic and per-round reflection surveys. We observe the fully multimodal model performs best in comparison to models trained on various compositions of modalities. We also find that the weights of some features (such as the expression of joy and the use of the word we) change in direction between the aforementioned models. We use these results to highlight the value of a multimodal dataset for studying competitive, collegiate debate.

📄 PDF Abstract BibTeX arXiv:2103.14189

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

End-to-End Audiovisual Fusion with LSTMs

2017-09-12 · Stavros Petridis, Yujiang Wang, Zuwei Li, Maja Pantic

Several end-to-end deep learning approaches have been recently presented which simultaneously extract visual features from the input images and perform visual speech classification. However, research on jointly extractin…

ClassificationGeneral Classificationspeech-recognitionSpeech Recognition

How Does Audio Influence Visual Attention in Omnidirectional Videos? Database and Model

2024-08-10 · Yuxin Zhu, Huiyu Duan, Kaiwei Zhang, Yucheng Zhu 외

Understanding and predicting viewer attention in omnidirectional videos (ODVs) is crucial for enhancing user engagement in virtual and augmented reality applications. Although both audio and visual modalities are essenti…

PredictionSaliency Prediction

Audio-Visual Quality Assessment for User Generated Content: Database and Method

2023-03-04 · Yuqin Cao, Xiongkuo Min, Wei Sun, XiaoPing Zhang 외

With the explosive increase of User Generated Content (UGC), UGC video quality assessment (VQA) becomes more and more important for improving users' Quality of Experience (QoE). However, most existing UGC VQA studies onl…

Video Quality AssessmentVisual Question Answering (VQA)

Subjective and Objective Audio-Visual Quality Assessment for User Generated Content

2023-07-10 · IEEE Transactions on Image Processing 2023 7 · Yuqin Cao, Xiongkuo Min, Wei Sun, Guangtao Zhai

In recent years, User Generated Content (UGC) has grown dramatically in video sharing applications. It is necessary for service-providers to use video quality assessment (VQA) to monitor and control users’ Quality of Exp…

Video Quality AssessmentVisual Question Answering (VQA)

Depression Scale Recognition from Audio, Visual and Text Analysis

2017-09-18 · Shubham Dham, Anirudh Sharma, Abhinav Dhall

Depression is a major mental health disorder that is rapidly affecting lives worldwide. Depression not only impacts emotional but also physical and psychological state of the person. Its symptoms include lack of interest…

Clustering