paper-with-me

홈 › Papers

Developing an Effective Training Dataset to Enhance the Performance of AI-based Speaker Separation Systems

2024-11-13 · Rawad Melhem, Assef Jafar, Oumayma Al Dakkak

This paper addresses the challenge of speaker separation, which remains an active research topic despite the promising results achieved in recent years. These results, however, often degrade in real recording conditions due to the presence of noise, echo, and other interferences. This is because neural models are typically trained on synthetic datasets consisting of mixed audio signals and their corresponding ground truths, which are generated using computer software and do not fully represent the complexities of real-world recording scenarios. The lack of realistic training sets for speaker separation remains a major hurdle, as obtaining individual sounds from mixed audio signals is a nontrivial task. To address this issue, we propose a novel method for constructing a realistic training set that includes mixture signals and corresponding ground truths for each speaker. We evaluate this dataset on a deep learning model and compare it to a synthetic dataset. We got a 1.65 dB improvement in Scale Invariant Signal to Distortion Ratio (SI-SDR) for speaker separation accuracy in realistic mixing. Our findings highlight the potential of realistic training sets for enhancing the performance of speaker separation models in real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2411.08375

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Separation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

HiLLIE: Human-in-the-Loop Training for Low-Light Image Enhancement

2025-05-04 · Xiaorui Zhao, Xinyue Zhou, Peibei Cao, Junyu Lou 외

Developing effective approaches to generate enhanced results that align well with human visual preferences for high-quality well-lit images remains a challenge in low-light image enhancement (LLIE). In this paper, we pro…

Image EnhancementImage Quality AssessmentLow-Light Image Enhancement

FINALLY: fast and universal speech enhancement with studio-like quality

2024-10-08 · Nicholas Babaev, Kirill Tamogashev, Azat Saginbaev, Ivan Shchekotov 외

In this paper, we address the challenge of speech enhancement in real-world recordings, which often contain various forms of distortion, such as background noise, reverberation, and microphone artifacts. We revisit the u…

Speech Enhancement

Dataset Regeneration for Sequential Recommendation

2024-05-28 · Mingjia Yin, Hao Wang, Wei Guo, Yong liu 외

The sequential recommender (SR) system is a crucial component of modern recommender systems, as it aims to capture the evolving preferences of users. Significant efforts have been made to enhance the capabilities of SR s…

Recommendation SystemsSequential Recommendation

Pretraining Strategies using Monolingual and Parallel Data for Low-Resource Machine Translation

2025-10-29 · Idriss Nguepi Nguefack, Mara Finkelstein, Toadoum Sari Sakayo arxiv

This research article examines the effectiveness of various pretraining strategies for developing machine translation models tailored to low-resource languages. Although this work considers several low-resource languages…

Machine Translation

UAV-Enhanced Combination to Application: Comprehensive Analysis and Benchmarking of a Human Detection Dataset for Disaster Scenarios

2024-08-09 · Ragib Amin Nihal, Benjamin Yen, Katsutoshi Itoyama, Kazuhiro Nakadai

Unmanned aerial vehicles (UAVs) have revolutionized search and rescue (SAR) operations, but the lack of specialized human detection datasets for training machine learning models poses a significant challenge.To address t…

BenchmarkingHuman DetectionObject Detection