paper-with-me

Papers

DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input

2022-08-22 · Jun Rekimoto

Interactions based on automatic speech recognition (ASR) have become widely used, with speech input being increasingly utilized to create documents. However, as there is no easy way to distinguish between commands being issued and text required to be input in speech, misrecognitions are difficult to identify and correct, meaning that documents need to be manually edited and corrected. The input of symbols and commands is also challenging because these may be misrecognized as text letters. To address these problems, this study proposes a speech interaction method called DualVoice, by which commands can be input in a whispered voice and letters in a normal voice. The proposed method does not require any specialized hardware other than a regular microphone, enabling a complete hands-free interaction. The method can be used in a wide range of situations where speech recognition is already available, ranging from text input to mobile/wearable computing. Two neural networks were designed in this study, one for discriminating normal speech from whispered speech, and the second for recognizing whisper speech. A prototype of a text input system was then developed to show how normal and whispered voice can be used in speech text input. Other potential applications using DualVoice are also discussed.

📄 PDF Abstract BibTeX arXiv:2208.10499

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Multi-modal Adversarial Training for Zero-Shot Voice Cloning

2024-08-28 · John Janiczek, Dading Chong, Dongyang Dai, Arlo Faria 외

A text-to-speech (TTS) model trained to reconstruct speech given text tends towards predictions that are close to the average characteristics of a dataset, failing to model the variations that make human speech sound nat…

Decodertext-to-speechText to SpeechVoice Cloning

Interpretable pap smear cell representation for cervical cancer screening

2023-11-17 · Yu Ando, Nora Jee-Young Park and, Gun Oh Chong, Seokhwan Ko 외

Screening is critical for prevention and early detection of cervical cancer but it is time-consuming and laborious. Supervised deep convolutional neural networks have been developed to automate pap smear screening and th…

ClusteringOne-Class Classification

The Speed-Vel Project: a Corpus of Acoustic and Aerodynamic Data to Measure Droplets Emission During Speech Interaction

2022-06-01 · LREC 2022 6 · Francesca Carbone, Gilles Bouchet, Alain Ghio, Thierry Legou 외

Conversations (normal speech) or professional interactions (e.g., projected speech in the classroom) have been identified as situations with increased risk of exposure to SARS-CoV-2 due to the high production of droplets…

Quartered Chirp Spectral Envelope for Whispered vs Normal Speech Classification

2024-08-27 · S. Johanan Joysingh, P. Vijayalakshmi, T. Nagarajan

Whispered speech as an acceptable form of human-computer interaction is gaining traction. Systems that address multiple modes of speech require a robust front-end speech classifier. Performance of whispered vs normal spe…

Exploiting Spectral Augmentation for Code-Switched Spoken Language Identification

2020-10-14 · Pradeep Rangan, Sundeep Teki, Hemant Misra

Spoken language Identification (LID) systems are needed to identify the language(s) present in a given audio sample, and typically could be the first step in many speech processing related tasks such as automatic speech …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+2