paper-with-me

홈 › Papers

Enhancing Audio Augmentation Methods with Consistency Learning

2021-02-09 · Turab Iqbal, Karim Helwani, Arvindh Krishnaswamy, Wenwu Wang

Data augmentation is an inexpensive way to increase training data diversity and is commonly achieved via transformations of existing data. For tasks such as classification, there is a good case for learning representations of the data that are invariant to such transformations, yet this is not explicitly enforced by classification losses such as the cross-entropy loss. This paper investigates the use of training objectives that explicitly impose this consistency constraint and how it can impact downstream audio classification tasks. In the context of deep convolutional neural networks in the supervised setting, we show empirically that certain measures of consistency are not implicitly captured by the cross-entropy loss and that incorporating such measures into the loss function can improve the performance of audio classification systems. Put another way, we demonstrate how existing augmentation methods can further improve learning by enforcing consistency.

📄 PDF Abstract BibTeX arXiv:2102.05151

Code (0)

등록된 구현이 없습니다.

Tasks

Audio ClassificationAudio TaggingClassificationData AugmentationDiversityGeneral Classification

Similar Papers 제목 키워드 기반

Improving Audio Event Recognition with Consistency Regularization

2025-09-12 · Shanmuka Sadhu, Weiran Wang arxiv

Consistency regularization (CR), which enforces agreement between model predictions on augmented views, has found recent benefits in automatic speech recognition [1]. In this paper, we propose the use of consistency regu…

Speech RecognitionData Augmentation

Exploring Self-Supervised Contrastive Learning of Spatial Sound Event Representation

2023-09-27 · Xilin Jiang, Cong Han, Yinghao Aaron Li, Nima Mesgarani

In this study, we present a simple multi-channel framework for contrastive learning (MC-SimCLR) to encode 'what' and 'where' of spatial audios. MC-SimCLR learns joint spectral and spatial representations from unlabeled s…

Contrastive LearningData Augmentation

AudioScenic: Audio-Driven Video Scene Editing

2024-04-25 · Kaixin Shen, Ruijie Quan, Linchao Zhu, Jun Xiao 외

Audio-driven visual scene editing endeavors to manipulate the visual background while leaving the foreground content unchanged, according to the given audio signals. Unlike current efforts focusing primarily on image edi…

CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional Video

2025-01-01 · CVPR 2025 1 · Zhaolin Wan, Han Qin, Zhiyang Li, Xiaopeng Fan 외

Omnidirectional videos (ODVs) present distinct challenges for accurate audio-visual saliency prediction due to their immersive nature, which combines spatial audio with panoramic visuals to enhance the user experienc…

Saliency Prediction

Video-to-Audio Generation with Hidden Alignment

2024-07-10 · Manjie Xu, Chenxing Li, Xinyi Tu, Yong Ren 외

Generating semantically and temporally aligned audio content in accordance with video input has become a focal point for researchers, particularly following the remarkable breakthrough in text-to-video generation. In thi…

Audio GenerationData AugmentationText-to-Video GenerationVideo Generation