paper-with-me

Papers

The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection

2024-09-17 · Gabriel Bibbó, Thomas Deacon, Arshdeep Singh, Mark D. Plumbley

This paper presents a residential audio dataset to support sound event detection research for smart home applications aimed at promoting wellbeing for older adults. The dataset is constructed by deploying audio recording systems in the homes of 8 participants aged 55-80 years for a 7-day period. Acoustic characteristics are documented through detailed floor plans and construction material information to enable replication of the recording environments for AI model deployment. A novel automated speech removal pipeline is developed, using pre-trained audio neural networks to detect and remove segments containing spoken voice, while preserving segments containing other sound events. The resulting dataset consists of privacy-compliant audio recordings that accurately capture the soundscapes and activities of daily living within residential spaces. The paper details the dataset creation methodology, the speech removal pipeline utilizing cascaded model architectures, and an analysis of the vocal label distribution to validate the speech removal process. This dataset enables the development and benchmarking of sound event detection models tailored specifically for in-home applications.

📄 PDF Abstract BibTeX arXiv:2409.11262

Code (1)

gbibbo/voice_anonymization 공식 구현 pytorch

Tasks

BenchmarkingEvent DetectionSound Event Detection

Similar Papers 제목 키워드 기반

Unified Audio Event Detection

2024-09-13 · Yidi Jiang, Ruijie Tao, Wen Huang, Qian Chen 외

Sound Event Detection (SED) detects regions of sound events, while Speaker Diarization (SD) segments speech conversations attributed to individual speakers. In SED, all speaker segments are classified as a single speech …

Event DetectionSound Event Detectionspeaker-diarizationSpeaker Diarization

Clotho: An Audio Captioning Dataset

2019-10-21 · Konstantinos Drossos, Samuel Lipping, Tuomas Virtanen

Audio captioning is the novel task of general audio content description using free text. It is an intermodal translation task (not speech-to-text), where a system accepts as an input an audio signal and outputs the textu…

Audio captioningDiversitySpeech-to-TextTranslation

Speech Denoising with Auditory Models

2020-11-21 · Mark R. Saddler, Andrew Francl, Jenelle Feather, Kaizhi Qian 외

Contemporary speech enhancement predominantly relies on audio transforms that are trained to reconstruct a clean speech waveform. The development of high-performing neural network sound recognition systems has raised the…

DenoisingSpeech DenoisingSpeech Enhancement

Learning Audio-Visual Dereverberation

2021-06-14 · Changan Chen, Wei Sun, David Harwath, Kristen Grauman

Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker IdentificationSpeech Enhancement+2

BioSEN: A Bio-acoustic Signal Enhancement Network for Animal Vocalizations

2026-05-02 · Tianyu Song, Ton Viet Ta, Ngamta Thamwattana, Hisako Nomura 외 arxiv

Most work in audio enhancement targets human speech, while bioacoustics is less studied due to noisy recordings and the distinct traits of animal sounds. To fill this gap, we adapt speech enhancement methods and build Bi…

Speech Enhancement