paper-with-me

Papers

3D Convolutional Neural Networks for Ultrasound-Based Silent Speech Interfaces

2021-04-23 · László Tóth, Amin Honarmandi Shandiz

Silent speech interfaces (SSI) aim to reconstruct the speech signal from a recording of the articulatory movement, such as an ultrasound video of the tongue. Currently, deep neural networks are the most successful technology for this task. The efficient solution requires methods that do not simply process single images, but are able to extract the tongue movement information from a sequence of video frames. One option for this is to apply recurrent neural structures such as the long short-term memory network (LSTM) in combination with 2D convolutional neural networks (CNNs). Here, we experiment with another approach that extends the CNN to perform 3D convolution, where the extra dimension corresponds to time. In particular, we apply the spatial and temporal convolutions in a decomposed form, which proved very successful recently in video action recognition. We find experimentally that our 3D network outperforms the CNN+LSTM model, indicating that 3D CNNs may be a feasible alternative to CNN+LSTM networks in SSI systems.

📄 PDF Abstract BibTeX arXiv:2104.11532

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionTemporal Action Localization

Methods 이 논문이 사용한 방법론

Memory Network 설명 없음

Similar Papers 제목 키워드 기반

SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

2026-08-01 · Ruidong Zhang, Jiacheng Liu, François Guimbretière, Cheng Zhang arxiv

Wearable silent speech interfaces (SSIs) are limited to small, closed vocabularies. Approaches achieving larger vocabularies require obtrusive hardware such as facial electrodes. We present SoniSpeech, the first large-sc…

Speech Recognition

Adaptation of Tongue Ultrasound-Based Silent Speech Interfaces Using Spatial Transformer Networks

2023-05-30 · László Tóth, Amin Honarmandi Shandiz, Gábor Gosztolya, Csapó Tamás Gábor

Thanks to the latest deep learning algorithms, silent speech interfaces (SSI) are now able to synthesize intelligible speech from articulatory movement data under certain conditions. However, the resulting models are rat…

Silent versus modal multi-speaker speech recognition from ultrasound and video

2021-02-27 · Manuel Sam Ribeiro, Aciel Eshky, Korin Richmond, Steve Renals

We investigate multi-speaker speech recognition from ultrasound images of the tongue and video images of the lips. We train our systems on imaging data from modal speech, and evaluate on matched test sets of two speaking…

Silent Speech Recognitionspeech-recognitionSpeech Recognition

Convolutional Neural Network-Based Age Estimation Using B-Mode Ultrasound Tongue Image

2021-01-27 · Kele Xu, Tamas Gábor Csapó, Ming Feng

Ultrasound tongue imaging is widely used for speech production research, and it has attracted increasing attention as its potential applications seem to be evident in many different fields, such as the visual biofeedback…

Age EstimationLanguage Acquisition

SottoVoce: An Ultrasound Imaging-Based Silent Speech Interaction Using Deep Neural Networks

2023-03-03 · Naoki Kimura, Michinari Kono, Jun Rekimoto

The availability of digital devices operated by voice is expanding rapidly. However, the applications of voice interfaces are still restricted. For example, speaking in public places becomes an annoyance to the surroundi…

speech-recognitionSpeech Recognition