paper-with-me

Papers

SSR7000: A Synchronized Corpus of Ultrasound Tongue Imaging for End-to-End Silent Speech Recognition

2022-06-01 · LREC 2022 6 · Naoki Kimura, Zixiong Su, Takaaki Saeki, Jun Rekimoto

This article presents SSR7000, a corpus of synchronized ultrasound tongue and lip images designed for end-to-end silent speech recognition (SSR). Although neural end-to-end models are successfully updating the state-of-the-art technology in the field of automatic speech recognition, SSR research based on ultrasound tongue imaging has still not evolved past cascaded DNN-HMM models due to the absence of a large dataset. In this study, we constructed a large dataset, namely SSR7000, to exploit the performance of the end-to-end models. The SSR7000 dataset contains ultrasound tongue and lip images of 7484 utterances by a single speaker. It contains more utterances per person than any other SSR corpus based on ultrasound imaging. We also describe preprocessing techniques to tackle data variances that are inevitable when collecting a large dataset and present benchmark results using an end-to-end model. The SSR7000 corpus is publicly available under the CC BY-NC 4.0 license.

📄 PDF Abstract BibTeX

Code (1)

supernaiter/ssr7000 공식 구현

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Silent Speech Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Silent versus modal multi-speaker speech recognition from ultrasound and video

2021-02-27 · Manuel Sam Ribeiro, Aciel Eshky, Korin Richmond, Steve Renals

We investigate multi-speaker speech recognition from ultrasound images of the tongue and video images of the lips. We train our systems on imaging data from modal speech, and evaluate on matched test sets of two speaking…

Silent Speech Recognitionspeech-recognitionSpeech Recognition

TaL: a synchronised multi-speaker corpus of ultrasound tongue imaging, audio, and lip videos

2020-11-19 · Manuel Sam Ribeiro, Jennifer Sanger, Jing-Xuan Zhang, Aciel Eshky 외

We present the Tongue and Lips corpus (TaL), a multi-speaker corpus of audio, ultrasound tongue imaging, and lip videos. TaL consists of two parts: TaL1 is a set of six recording sessions of one professional voice talent…

speech-recognitionSpeech RecognitionSpeech Synthesis

Convolutional Neural Network-Based Age Estimation Using B-Mode Ultrasound Tongue Image

2021-01-27 · Kele Xu, Tamas Gábor Csapó, Ming Feng

Ultrasound tongue imaging is widely used for speech production research, and it has attracted increasing attention as its potential applications seem to be evident in many different fields, such as the visual biofeedback…

Age EstimationLanguage Acquisition

Adaptation of Tongue Ultrasound-Based Silent Speech Interfaces Using Spatial Transformer Networks

2023-05-30 · László Tóth, Amin Honarmandi Shandiz, Gábor Gosztolya, Csapó Tamás Gábor

Thanks to the latest deep learning algorithms, silent speech interfaces (SSI) are now able to synthesize intelligible speech from articulatory movement data under certain conditions. However, the resulting models are rat…

Contour-based 3d tongue motion visualization using ultrasound image sequences

2016-05-19 · Kele Xu, Yin Yang, Clémence Leboullenger, Pierre Roussel 외

This article describes a contour-based 3D tongue deformation visualization framework using B-mode ultrasound image sequences. A robust, automatic tracking algorithm characterizes tongue motion via a contour, which is the…

Silent Speech Recognitionspeech-recognitionSpeech Recognition