CirdoX: an on/off-line multisource speech and sound analysis software
Vocal User Interfaces in domestic environments recently gained interest in the speech processing community. This interest is due to the opportunity of using it in the framework of Ambient Assisted Living both for home automation (vocal command) and for call for help in case of distress situations, i.e. after a fall. C IRDO X, which is a modular software, is able to analyse online the audio environment in a home, to extract the uttered sentences and then to process them thanks to an ASR module. Moreover, this system perfoms non-speech audio event classification; in this case, specific models must be trained. The software is designed to be modular and to process on-line the audio multichannel stream. Some exemples of studies in which C IRDO X was involved are described. They were operated in real environment, namely a Living lab environment.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationSimilar Papers 제목 키워드 기반
Reconnaissance automatique de la parole distante dans un habitat intelligent : m\'ethodes multi-sources en conditions r\'ealistes (Distant Speech Recognition in a Smart Home : Comparison of Several Multisource ASRs in Realistic Conditions) [in French]
Indian EmoSpeech Command Dataset: A dataset for emotion based speech recognition in the wild
Speech emotion analysis is an important task which further enables several application use cases. The non-verbal sounds within speech utterances also play a pivotal role in emotion analysis in speech. Due to the widespre…
Emotion RecognitionKeyword Spottingspeech-recognitionSpeech RecognitionTTMBA: Towards Text To Multiple Sources Binaural Audio Generation
Most existing text-to-audio (TTA) generation methods produce mono outputs, neglecting essential spatial information for immersive auditory experiences. To address this issue, we propose a cascaded method for text-to-mult…
Audio GenerationUnified Audio Event Detection
Sound Event Detection (SED) detects regions of sound events, while Speaker Diarization (SD) segments speech conversations attributed to individual speakers. In SED, all speaker segments are classified as a single speech …
Event DetectionSound Event Detectionspeaker-diarizationSpeaker DiarizationModeling and Measuring Redundancy in Multisource Multimodal Data for Autonomous Driving
Next-generation autonomous vehicles (AVs) rely on large volumes of multisource and multimodal ($M^2$) data to support real-time decision-making. In practice, data quality (DQ) varies across sources and modalities due to …
Autonomous VehiclesAutonomous DrivingObject Detection