paper-with-me

홈 › Papers

Improving Audio Event Recognition with Consistency Regularization

2025-09-12 · Shanmuka Sadhu, Weiran Wang arxiv

Consistency regularization (CR), which enforces agreement between model predictions on augmented views, has found recent benefits in automatic speech recognition [1]. In this paper, we propose the use of consistency regularization for audio event recognition, and demonstrate its effectiveness on AudioSet. With extensive ablation studies for both small ($\sim$20k) and large ($\sim$1.8M) supervised training sets, we show that CR brings consistent improvement over supervised baselines which already heavily utilize data augmentation, and CR using stronger augmentation and multiple augmentations leads to additional gain for the small training set. Furthermore, we extend the use of CR into the semi-supervised setup with 20K labeled samples and 1.8M unlabeled samples, and obtain performance improvement over our best model trained on the small set.

📄 PDF Abstract BibTeX arXiv:2509.10391

Code (0)

등록된 구현이 없습니다.

Tasks

Speech RecognitionData Augmentation

Similar Papers 제목 키워드 기반

Hierarchical Semantic-Constrained Heterogeneous Graph for Audio-Visual Event Localization

2026-06-05 · Zhe Yang, Ruyi Zhang, Hongtao Chen, Wenrui Li 외 arxiv

Open-vocabulary audio-visual event localization (OV-AVEL) jointly models audio-visual cues to recognize and temporally localize events, including categories unseen during training. Existing methods primarily learn joint …

audio-visual event localization

Semi-supervised Sound Event Detection with Local and Global Consistency Regularization

2023-09-15 · Yiming Li, Xiangdong Wang, Hong Liu, Rui Tao 외

Learning meaningful frame-wise features on a partially labeled dataset is crucial to semi-supervised sound event detection. Prior works either maintain consistency on frame-level predictions or seek feature-level similar…

Event DetectionSound Event Detection

Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization

2025-10-27 · Bernardo Torres, Manuel Moussallam, Gabriel Meseguer-Brocal arxiv

Audio autoencoders learn useful, compressed audio representations, but their non-linear latent spaces prevent intuitive algebraic manipulation such as mixing or scaling. We introduce a simple training methodology to indu…

Data Augmentation

ROCM: RLHF on consistency models

2025-03-08 · Shivanshu Shekhar, Tong Zhang

Diffusion models have revolutionized generative modeling in continuous domains like image, audio, and video synthesis. However, their iterative sampling process leads to slow generation and inefficient training, challeng…

Policy Gradient Methods

Towards Open-Vocabulary Audio-Visual Event Localization

2024-11-18 · CVPR 2025 1 · Jinxing Zhou, Dan Guo, Ruohao Guo, Yuxin Mao 외

The Audio-Visual Event Localization (AVEL) task aims to temporally locate and classify video events that are both audible and visible. Most research in this field assumes a closed-set setting, which restricts these model…

audio-visual event localization