paper-with-me

홈 › Papers

Ambient Sound Helps: Audiovisual Crowd Counting in Extreme Conditions

2020-05-14 · Di Hu, Lichao Mou, Qingzhong Wang, Junyu. Gao, Yuansheng Hua, Dejing Dou, Xiao Xiang Zhu

Visual crowd counting has been recently studied as a way to enable people counting in crowd scenes from images. Albeit successful, vision-based crowd counting approaches could fail to capture informative features in extreme conditions, e.g., imaging at night and occlusion. In this work, we introduce a novel task of audiovisual crowd counting, in which visual and auditory information are integrated for counting purposes. We collect a large-scale benchmark, named auDiovISual Crowd cOunting (DISCO) dataset, consisting of 1,935 images and the corresponding audio clips, and 170,270 annotated instances. In order to fuse the two modalities, we make use of a linear feature-wise fusion module that carries out an affine transformation on visual and auditory features. Finally, we conduct extensive experiments using the proposed dataset and approach. Experimental results show that introducing auditory information can benefit crowd counting under different illumination, noise, and occlusion conditions. The dataset and code will be released. Code and data have been made available

📄 PDF Abstract BibTeX arXiv:2005.07097

Code (1)

qingzwang/AudioVisualCrowdCounting 공식 구현 pytorch

Tasks

Crowd Counting

Similar Papers 제목 키워드 기반

Learning What To Hear: Boosting Sound-Source Association For Robust Audiovisual Instance Segmentation

2025-09-26 · Jinbae Seo, Hyeongjun Kwon, Kwonyoung Kim, Jiyoung Lee 외 arxiv

Audiovisual instance segmentation (AVIS) requires accurately localizing and tracking sounding objects throughout video sequences. Existing methods suffer from visual bias stemming from two fundamental issues: uniform add…

Instance Segmentation

Repetitive Activity Counting by Sight and Sound

2021-03-24 · CVPR 2021 1 · Yunhua Zhang, Ling Shao, Cees G. M. Snoek

This paper strives for repetitive activity counting in videos. Different from existing works, which all analyze the visual video content only, we incorporate for the first time the corresponding sound into the repetition…

Looking Similar Sounding Different: Leveraging Counterfactual Cross-Modal Pairs for Audiovisual Representation Learning

2024-01-01 · CVPR 2024 1 · Nikhil Singh, Chih-Wei Wu, Iroro Orife, Mahdi Kalayeh

Audiovisual representation learning typically relies on the correspondence between sight and sound. However there are often multiple audio tracks that can correspond with a visual scene. Consider for example differen…

Contrastive LearningcounterfactualRepresentation Learning

Looking Similar, Sounding Different: Leveraging Counterfactual Cross-Modal Pairs for Audiovisual Representation Learning

2023-04-12 · Nikhil Singh, Chih-Wei Wu, Iroro Orife, Mahdi Kalayeh

Audiovisual representation learning typically relies on the correspondence between sight and sound. However, there are often multiple audio tracks that can correspond with a visual scene. Consider, for example, different…

Contrastive LearningcounterfactualRepresentation Learning

ESSYS* Sharing #UC: An Emotion-driven Audiovisual Installation

2022-09-07 · Sérgio M. Rebelo, Mariana Seiça, Pedro Martins, João Bicker 외

We present ESSYS* Sharing #UC, an audiovisual installation artwork that reflects upon the emotional context related to the university and the city of Coimbra, based on the data shared about them on Twitter. The installat…

Diversity