Addressing Feature Imbalance in Sound Source Separation
Neural networks often suffer from a feature preference problem, where they tend to overly rely on specific features to solve a task while disregarding other features, even if those neglected features are essential for the task. Feature preference problems have primarily been investigated in classification task. However, we observe that feature preference occurs in high-dimensional regression task, specifically, source separation. To mitigate feature preference in source separation, we propose FEAture BAlancing by Suppressing Easy feature (FEABASE). This approach enables efficient data utilization by learning hidden information about the neglected feature. We evaluate our method in a multi-channel source separation task, where feature preference between spatial feature and timbre feature appears.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
This paper presents a framework for universal sound separation and polyphonic audio classification, addressing the challenges of separating and classifying individual sound sources in a multichannel mixture. The proposed…
Audio ClassificationClassificationMambaSemantic Grouping Network for Audio Source Separation
Recently, audio-visual separation approaches have taken advantage of the natural synchronization between the two modalities to boost audio source separation performance. They extracted high-level semantics from visual in…
Audio Source SeparationLeveraging Sound Source Trajectories for Universal Sound Separation
Existing methods utilizing spatial information for sound source separation require prior knowledge of the direction of arrival (DOA) of the source or utilize estimated but imprecise localization results, which impairs th…
Sound Source LocalizationSelf-Supervised Audio-Visual Co-Segmentation
Segmenting objects in images and separating sound sources in audio are challenging tasks, in part because traditional approaches require large amounts of labeled data. In this paper we develop a neural network model for …
Image SegmentationSegmentationSemantic SegmentationVisually-Guided Sound Source Separation with Audio-Visual Predictive Coding
The framework of visually-guided sound source separation generally consists of three parts: visual feature extraction, multimodal feature fusion, and sound signal processing. An ongoing trend in this field has been to ta…
validVisually Guided Sound Source Separation