The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
This paper presents a residential audio dataset to support sound event detection research for smart home applications aimed at promoting wellbeing for older adults. The dataset is constructed by deploying audio recording systems in the homes of 8 participants aged 55-80 years for a 7-day period. Acoustic characteristics are documented through detailed floor plans and construction material information to enable replication of the recording environments for AI model deployment. A novel automated speech removal pipeline is developed, using pre-trained audio neural networks to detect and remove segments containing spoken voice, while preserving segments containing other sound events. The resulting dataset consists of privacy-compliant audio recordings that accurately capture the soundscapes and activities of daily living within residential spaces. The paper details the dataset creation methodology, the speech removal pipeline utilizing cascaded model architectures, and an analysis of the vocal label distribution to validate the speech removal process. This dataset enables the development and benchmarking of sound event detection models tailored specifically for in-home applications.
Code (1)
Tasks
BenchmarkingEvent DetectionSound Event DetectionSimilar Papers 제목 키워드 기반
Unified Audio Event Detection
Sound Event Detection (SED) detects regions of sound events, while Speaker Diarization (SD) segments speech conversations attributed to individual speakers. In SED, all speaker segments are classified as a single speech …
Event DetectionSound Event Detectionspeaker-diarizationSpeaker DiarizationClotho: An Audio Captioning Dataset
Audio captioning is the novel task of general audio content description using free text. It is an intermodal translation task (not speech-to-text), where a system accepts as an input an audio signal and outputs the textu…
Audio captioningDiversitySpeech-to-TextTranslationSpeech Denoising with Auditory Models
Contemporary speech enhancement predominantly relies on audio transforms that are trained to reconstruct a clean speech waveform. The development of high-performing neural network sound recognition systems has raised the…
DenoisingSpeech DenoisingSpeech EnhancementLearning Audio-Visual Dereverberation
Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker IdentificationSpeech Enhancement+2BioSEN: A Bio-acoustic Signal Enhancement Network for Animal Vocalizations
Most work in audio enhancement targets human speech, while bioacoustics is less studied due to noisy recordings and the distinct traits of animal sounds. To fill this gap, we adapt speech enhancement methods and build Bi…
Speech Enhancement