paper-with-me

Papers

Data centric approach to Chinese Medical Speech Recognition

2021-10-01 · ROCLING 2021 10 · Sheng-Luen Chung, Yi-Shiuan Li, Hsien-Wei Ting

Concerning the development of Chinese medical speech recognition technology, this study re-addresses earlier encountered issues in accordance with the process of Machine Learning Engineering for Production (MLOps) from a data centric perspective. First is the new segmentation of speech utterances to meet sentences completeness for all utterances in the collected Chinese Medical Speech Corpus (ChiMeS). Second is optimization of Joint CTC/Attention model through data augmentation in boosting recognition performance out of very limited speech corpus. Overall, to facilitate the development of Chinese medical speech recognition, this paper contributes: (1) The ChiMeS corpus, the first Chinese Medicine Speech corpus of its kind, which is 14.4 hours, with a total of 7,225 sentences. (2) A trained Joint CTC/Attention ASR model by ChiMeS-14, yielding a Character Error Rate (CER) of 13.65% and a Keyword Error Rate (KER) of 20.82%, respectively, when tested on the ChiMeS-14 testing set. And (3) an evaluation platform set up to compare performance of other ASR models. All the released resources can be found in the ChiMeS portal (https://iclab.ee.ntust.edu.tw/home).

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Towards Multilingual Conversations in the Medical Domain: Development of Multilingual Medical Data and A Network-based ASR System

2014-05-01 · LREC 2014 5 · Sakriani Sakti, Keigo Kubo, Sho Matsumiya, Graham Neubig 외

This paper outlines the recent development on multilingual medical data and multilingual speech recognition system for network-based speech-to-speech translation in the medical domain. The overall speech-to-speech transl…

Machine Translationspeech-recognitionSpeech RecognitionSpeech Synthesis+3

Chinese Medical Speech Recognition with Punctuated Hypothesis

2021-10-01 · ROCLING 2021 10 · Sheng-Luen Chung, Jin-Huan Fan, Hsien-Wei Ting

Automatic Speech Recognition (ASR) technology presents the possibility for medical professionals to document patient record, diagnosis, postoperative care, patrol records, and etc. that are now done manually. However, ea…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Speech Recognition With No Speech Or With Noisy Speech Beyond English

2019-06-17 · Gautam Krishna, Co Tran, Yan Han, Mason Carnahan 외

In this paper we demonstrate continuous noisy speech recognition using connectionist temporal classification (CTC) model on limited Chinese vocabulary using electroencephalography (EEG) features with no speech signal as …

EEGElectroencephalogram (EEG)General ClassificationNoisy Speech Recognition+2

Exploring Word Segmentation and Medical Concept Recognition for Chinese Medical Texts

2021-06-01 · NAACL (BioNLP) 2021 6 · Yang Liu, Yuanhe Tian, Tsung-Hui Chang, Song Wu 외

Chinese word segmentation (CWS) and medical concept recognition are two fundamental tasks to process Chinese electronic medical records (EMRs) and play important roles in downstream tasks for understanding Chinese EMRs. …

Chinese Word SegmentationModel SelectionSegmentation

A Small and Fast BERT for Chinese Medical Punctuation Restoration

2023-08-24 · Tongtao Ling, Yutao Lai, Lei Chen, Shilei Huang 외

In clinical dictation, utterances after automatic speech recognition (ASR) without explicit punctuation marks may lead to the misunderstanding of dictated reports. To give a precise and understandable clinical report wit…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive LearningPunctuation Restoration+2