paper-with-me

Papers

Speech-dependent Data Augmentation for Own Voice Reconstruction with Hearable Microphones in Noisy Environments

2024-05-19 · Mattes Ohlenbusch, Christian Rollwage, Simon Doclo

Own voice pickup for hearables in noisy environments benefits from using both an outer and an in-ear microphone outside and inside the occluded ear. Due to environmental noise recorded at both microphones, and amplification of the own voice at low frequencies and band-limitation at the in-ear microphone, an own voice reconstruction system is needed to enable communication. A large amount of own voice signals is required to train a supervised deep learning-based own voice reconstruction system. Training data can either be obtained by recording a large amount of own voice signals of different talkers with a specific device, which is costly, or through augmentation of available speech data. Own voice signals can be simulated by assuming a linear time-invariant relative transfer function between hearable microphones for each phoneme, referred to as own voice transfer characteristics. In this paper, we propose data augmentation techniques for training an own voice reconstruction system based on speech-dependent models of own voice transfer characteristics between hearable microphones. The proposed techniques use few recorded own voice signals to estimate transfer characteristics and can then be used to simulate a large amount of own voice signals based on single-channel speech signals. Experimental results show that the proposed speech-dependent individual data augmentation technique leads to better performance compared to other data augmentation techniques or compared to training only on the available recorded own voice signals, and additional fine-tuning on the available recorded signals can improve performance further.

📄 PDF Abstract BibTeX arXiv:2405.11592

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

Low-Complexity Own Voice Reconstruction for Hearables with an In-Ear Microphone

2024-09-06 · Mattes Ohlenbusch, Christian Rollwage, Simon Doclo

Hearable devices, equipped with one or more microphones, are commonly used for speech communication. Here, we consider the scenario where a hearable is used to capture the user's own voice in a noisy environment. In this…

Data Augmentation

Multi-Microphone Noise Data Augmentation for DNN-based Own Voice Reconstruction for Hearables in Noisy Environments

2023-12-14 · Mattes Ohlenbusch, Christian Rollwage, Simon Doclo

Hearables with integrated microphones may offer communication benefits in noisy working environments, e.g. by transmitting the recorded own voice of the user. Systems aiming at reconstructing the clean and full-bandwidth…

Data Augmentation

An Improved StarGAN for Emotional Voice Conversion: Enhancing Voice Quality and Data Augmentation

2021-07-18 · Xiangheng He, Junjie Chen, Georgios Rizos, Björn W. Schuller

Emotional Voice Conversion (EVC) aims to convert the emotional style of a source speech signal to a target style while preserving its content and speaker identity information. Previous emotional conversion studies do not…

Data AugmentationEmotion RecognitionGenerative Adversarial NetworkSpeech Emotion Recognition+1

Speech Reconstruction with Reminiscent Sound via Visual Voice Memory

2021-11-17 · IEEE/ACM Transactions on Audio, Speech, and Language Processing 2021 11 · Joanna Hong, Minsu Kim, Se Jin Park, Yong Man Ro

The goal of this work is to reconstruct speech from silent video, in both speaker dependent and independent ways. Unlike previous works that have been mostly restricted to a speaker dependent setting, we propose Visual V…

Speaker-Specific Lip to Speech Synthesis

Non-Parallel Voice Conversion for ASR Augmentation

2022-09-15 · Gary Wang, Andrew Rosenberg, Bhuvana Ramabhadran, Fadi Biadsy 외

Automatic speech recognition (ASR) needs to be robust to speaker differences. Voice Conversion (VC) modifies speaker characteristics of input speech. This is an attractive feature for ASR data augmentation. In this paper…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDiversity+3