paper-with-me

Papers

Adaptation of Tongue Ultrasound-Based Silent Speech Interfaces Using Spatial Transformer Networks

2023-05-30 · László Tóth, Amin Honarmandi Shandiz, Gábor Gosztolya, Csapó Tamás Gábor

Thanks to the latest deep learning algorithms, silent speech interfaces (SSI) are now able to synthesize intelligible speech from articulatory movement data under certain conditions. However, the resulting models are rather speaker-specific, making a quick switch between users troublesome. Even for the same speaker, these models perform poorly cross-session, i.e. after dismounting and re-mounting the recording equipment. To aid quick speaker and session adaptation of ultrasound tongue imaging-based SSI models, we extend our deep networks with a spatial transformer network (STN) module, capable of performing an affine transformation on the input images. Although the STN part takes up only about 10% of the network, our experiments show that adapting just the STN module might allow to reduce MSE by 88% on the average, compared to retraining the whole network. The improvement is even larger (around 92%) when adapting the network to different recording sessions from the same speaker.

📄 PDF Abstract BibTeX arXiv:2305.19130

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Spatial Transformer A Spatial Transformer is an image model block that explicitly allows the spatial manipulation of data within a [convolutional neural…

Similar Papers 제목 키워드 기반

Silent versus modal multi-speaker speech recognition from ultrasound and video

2021-02-27 · Manuel Sam Ribeiro, Aciel Eshky, Korin Richmond, Steve Renals

We investigate multi-speaker speech recognition from ultrasound images of the tongue and video images of the lips. We train our systems on imaging data from modal speech, and evaluate on matched test sets of two speaking…

Silent Speech Recognitionspeech-recognitionSpeech Recognition

3D Convolutional Neural Networks for Ultrasound-Based Silent Speech Interfaces

2021-04-23 · László Tóth, Amin Honarmandi Shandiz

Silent speech interfaces (SSI) aim to reconstruct the speech signal from a recording of the articulatory movement, such as an ultrasound video of the tongue. Currently, deep neural networks are the most successful techno…

Action RecognitionTemporal Action Localization

SSR7000: A Synchronized Corpus of Ultrasound Tongue Imaging for End-to-End Silent Speech Recognition

2022-06-01 · LREC 2022 6 · Naoki Kimura, Zixiong Su, Takaaki Saeki, Jun Rekimoto

This article presents SSR7000, a corpus of synchronized ultrasound tongue and lip images designed for end-to-end silent speech recognition (SSR). Although neural end-to-end models are successfully updating the state-of-t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Silent Speech Recognitionspeech-recognition+1

Contour-based 3d tongue motion visualization using ultrasound image sequences

2016-05-19 · Kele Xu, Yin Yang, Clémence Leboullenger, Pierre Roussel 외

This article describes a contour-based 3D tongue deformation visualization framework using B-mode ultrasound image sequences. A robust, automatic tracking algorithm characterizes tongue motion via a contour, which is the…

Silent Speech Recognitionspeech-recognitionSpeech Recognition

Movement Detection of Tongue and Related Body Parts Using IR-UWB Radar

2022-09-05 · Sunghwa Lee, Younghoon Shin

Because an impulse radio ultra-wideband (IR-UWB) radar can detect targets with high accuracy, work through occluding materials, and operate without contact, it is an attractive hardware solution for building silent speec…