paper-with-me

Papers

An Audio-Visual Dataset and Deep Learning Frameworks for Crowded Scene Classification

2021-12-16 · Lam Pham, Dat Ngo, Phu X. Nguyen, Truong Hoang, Alexander Schindler

This paper presents a task of audio-visual scene classification (SC) where input videos are classified into one of five real-life crowded scenes: 'Riot', 'Noise-Street', 'Firework-Event', 'Music-Event', and 'Sport-Atmosphere'. To this end, we firstly collect an audio-visual dataset (videos) of these five crowded contexts from Youtube (in-the-wild scenes). Then, a wide range of deep learning frameworks are proposed to deploy either audio or visual input data independently. Finally, results obtained from high-performed deep learning frameworks are fused to achieve the best accuracy score. Our experimental results indicate that audio and visual input factors independently contribute to the SC task's performance. Significantly, an ensemble of deep learning frameworks exploring either audio or visual input data can achieve the best accuracy of 95.7%.

📄 PDF Abstract BibTeX arXiv:2112.09172

Code (1)

phamdanglam1986/An-application-demo-of-audio-visual-crowded-scene-classification- 공식 구현

Tasks

Deep LearningScene Classification

Similar Papers 제목 키워드 기반

No-audio speaking status detection in crowded settings via visual pose-based filtering and wearable acceleration

2022-11-01 · Jose Vargas-Quiros, Laura Cabrera-Quiros, Hayley Hung

Recognizing who is speaking in a crowded scene is a key challenge towards the understanding of the social interactions going on within. Detecting speaking status from body movement alone opens the door for the analysis o…

Action RecognitionPrivacy Preserving

Looking Similar Sounding Different: Leveraging Counterfactual Cross-Modal Pairs for Audiovisual Representation Learning

2024-01-01 · CVPR 2024 1 · Nikhil Singh, Chih-Wei Wu, Iroro Orife, Mahdi Kalayeh

Audiovisual representation learning typically relies on the correspondence between sight and sound. However there are often multiple audio tracks that can correspond with a visual scene. Consider for example differen…

Contrastive LearningcounterfactualRepresentation Learning

Looking Similar, Sounding Different: Leveraging Counterfactual Cross-Modal Pairs for Audiovisual Representation Learning

2023-04-12 · Nikhil Singh, Chih-Wei Wu, Iroro Orife, Mahdi Kalayeh

Audiovisual representation learning typically relies on the correspondence between sight and sound. However, there are often multiple audio tracks that can correspond with a visual scene. Consider, for example, different…

Contrastive LearningcounterfactualRepresentation Learning

Crowded Scene Analysis: A Survey

2015-02-06 · Teng Li, Huan Chang, Meng Wang, Bingbing Ni 외

Automated scene analysis has been a topic of great interest in computer vision and cognitive science. Recently, with the growth of crowd phenomena in the real world, crowded scene analysis has attracted much attention. H…

Anomaly DetectionSurvey

Detection in Crowded Scenes: One Proposal, Multiple Predictions

2020-03-20 · CVPR 2020 6 · Xuangeng Chu, Anlin Zheng, Xiangyu Zhang, Jian Sun

We propose a simple yet effective proposal-based object detector, aiming at detecting highly-overlapped instances in crowded scenes. The key of our approach is to let each proposal predict a set of correlated instances r…

Object DetectionPedestrian Detection