paper-with-me

홈 › Papers

AVCAffe: A Large Scale Audio-Visual Dataset of Cognitive Load and Affect for Remote Work

2022-05-13 · Pritam Sarkar, Aaron Posen, Ali Etemad

We introduce AVCAffe, the first Audio-Visual dataset consisting of Cognitive load and Affect attributes. We record AVCAffe by simulating remote work scenarios over a video-conferencing platform, where subjects collaborate to complete a number of cognitively engaging tasks. AVCAffe is the largest originally collected (not collected from the Internet) affective dataset in English language. We recruit 106 participants from 18 different countries of origin, spanning an age range of 18 to 57 years old, with a balanced male-female ratio. AVCAffe comprises a total of 108 hours of video, equivalent to more than 58,000 clips along with task-based self-reported ground truth labels for arousal, valence, and cognitive load attributes such as mental demand, temporal demand, effort, and a few others. We believe AVCAffe would be a challenging benchmark for the deep learning research community given the inherent difficulty of classifying affect and cognitive load in particular. Moreover, our dataset fills an existing timely gap by facilitating the creation of learning systems for better self-management of remote work meetings, and further study of hypotheses regarding the impact of remote work on cognitive load and affective states.

📄 PDF Abstract BibTeX arXiv:2205.06887

Code (1)

pritamqu/AVCAffe 공식 구현 pytorch

Tasks

Management

Similar Papers 제목 키워드 기반

M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment

2024-03-14 · Long Nguyen-Phuoc, Renald Gaboriau, Dimitri Delacroix, Laurent Navarro

This paper introduces the M&M model, a novel multimodal-multitask learning framework, applied to the AVCAffe dataset for cognitive load assessment (CLA). M&M uniquely integrates audiovisual cues through a dual-pathway ar…

VGGSound: A Large-scale Audio-Visual Dataset

2020-04-29 · Honglie Chen, Weidi Xie, Andrea Vedaldi, Andrew Zisserman

Our goal is to collect a large-scale audio-visual dataset with low label noise from videos in the wild using computer vision techniques. The resulting dataset can be used for training and evaluating audio recognition mod…

image-classificationImage Classification

Audiovisual Moments in Time: A Large-Scale Annotated Dataset of Audiovisual Actions

2023-08-18 · PLOS ONE 2024 4 · Michael Joannou, Pia Rotshtein, Uta Noppeney

We present Audiovisual Moments in Time (AVMIT), a large-scale dataset of audiovisual action events. In an extensive annotation task 11 participants labelled a subset of 3-second audiovisual videos from the Moments in Tim…

Aligned Better, Listen Better for Audio-Visual Large Language Models

2025-04-02 · Yuxin Guo, Shuailei Ma, Shijie Ma, Xiaoyi Bao 외

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter…

Video Understanding

Speaker Recognition in Realistic Scenario Using Multimodal Data

2023-02-25 · Saqlain Hussain Shah, Muhammad Saad Saeed, Shah Nawaz, Muhammad Haroon Yousaf

In recent years, an association is established between faces and voices of celebrities leveraging large scale audio-visual information from YouTube. The availability of large scale audio-visual datasets is instrumental i…

Speaker Recognition