paper-with-me

홈 › Papers

Deep Learning for Punctuation Restoration in Medical Reports

2017-08-01 · WS 2017 8 · Wael Salloum, Greg Finley, Erik Edwards, Mark Miller, David Suendermann-Oeft

In clinical dictation, speakers try to be as concise as possible to save time, often resulting in utterances without explicit punctuation commands. Since the end product of a dictated report, e.g. an out-patient letter, does require correct orthography, including exact punctuation, the latter need to be restored, preferably by automated means. This paper describes a method for punctuation restoration based on a state-of-the-art stack of NLP and machine learning techniques including B-RNNs with an attention mechanism and late fusion, as well as a feature extraction technique tailored to the processing of medical terminology using a novel vocabulary reduction model. To the best of our knowledge, the resulting performance is superior to that reported in prior art on similar tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningPunctuation RestorationSpeech Recognition

Similar Papers 제목 키워드 기반

A Small and Fast BERT for Chinese Medical Punctuation Restoration

2023-08-24 · Tongtao Ling, Yutao Lai, Lei Chen, Shilei Huang 외

In clinical dictation, utterances after automatic speech recognition (ASR) without explicit punctuation marks may lead to the misunderstanding of dictated reports. To give a precise and understandable clinical report wit…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive LearningPunctuation Restoration+2

QASR: QCRI Aljazeera Speech Resource A Large Scale Annotated Arabic Speech Corpus

2021-08-01 · ACL 2021 5 · Hamdy Mubarak, Amir Hussein, Shammur Absar Chowdhury, Ahmed Ali

We introduce the largest transcribed Arabic speech corpus, QASR, collected from the broadcast domain. This multi-dialect speech dataset contains 2,000 hours of speech sampled at 16kHz crawled from Aljazeera news channel.…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Dialect IdentificationLanguage Modeling+8

QASR: QCRI Aljazeera Speech Resource -- A Large Scale Annotated Arabic Speech Corpus

2021-06-24 · Hamdy Mubarak, Amir Hussein, Shammur Absar Chowdhury, Ahmed Ali

We introduce the largest transcribed Arabic speech corpus, QASR, collected from the broadcast domain. This multi-dialect speech dataset contains 2,000 hours of speech sampled at 16kHz crawled from Aljazeera news channel.…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Dialect IdentificationLanguage Modeling+8

PersianPunc: A Large-Scale Dataset and BERT-Based Approach for Persian Punctuation Restoration

2026-03-05 · Mohammad Javad Ranjbar Kalahroodi, Heshaam Faili, Azadeh Shakery arxiv

Punctuation restoration is essential for improving the readability and downstream utility of automatic speech recognition (ASR) outputs, yet remains underexplored for Persian despite its importance. We introduce PersianP…

Speech Recognition

Token-Level Supervised Contrastive Learning for Punctuation Restoration

2021-07-19 · Qiushi Huang, Tom Ko, H Lilian Tang, Xubo Liu 외

Punctuation is critical in understanding natural language text. Currently, most automatic speech recognition (ASR) systems do not generate punctuation, which affects the performance of downstream tasks, such as intent de…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive LearningIntent Detection+5