paper-with-me

Papers

A Study into Pre-training Strategies for Spoken Language Understanding on Dysarthric Speech

2021-06-15 · Pu Wang, Bagher BabaAli, Hugo Van hamme

End-to-end (E2E) spoken language understanding (SLU) systems avoid an intermediate textual representation by mapping speech directly into intents with slot values. This approach requires considerable domain-specific training data. In low-resource scenarios this is a major concern, e.g., in the present study dealing with SLU for dysarthric speech. Pretraining part of the SLU model for automatic speech recognition targets helps but no research has shown to which extent SLU on dysarthric speech benefits from knowledge transferred from other dysarthric speech tasks. This paper investigates the efficiency of pre-training strategies for SLU tasks on dysarthric speech. The designed SLU system consists of a TDNN acoustic model for feature encoding and a capsule network for intent and slot decoding. The acoustic model is pre-trained in two stages: initialization with a corpus of normal speech and finetuning on a mixture of dysarthric and normal speech. By introducing the intelligibility score as a metric of the impairment severity, this paper quantitatively analyzes the relation between generalization and pathology severity for dysarthric speech.

📄 PDF Abstract BibTeX arXiv:2106.08313

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSpoken Language Understanding

Methods 이 논문이 사용한 방법론

Capsule Network A capsule is an activation vector that basically executes on its inputs some complex internal computations. Length of these activation vectors signifies the probability of…

Similar Papers 제목 키워드 기반

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue

2026-05-11 · Hui Lu, Xueyuan Chen, Huimeng Wang, Shuhai Peng 외 arxiv

Full-duplex spoken dialogue requires a model to keep listening while generating its own spoken response. This is challenging for large language models (LLMs), which are designed to extend a single coherent sequence and d…

Question Answering

Improved Long-Form Spoken Language Translation with Large Language Models

2022-12-19 · Arya D. McCarthy, Hao Zhang, Shankar Kumar, Felix Stahlberg 외

A challenge in spoken language translation is that plenty of spoken content is long-form, but short units are necessary for obtaining high-quality translations. To address this mismatch, we fine-tune a general-purpose, l…

FormLanguage ModelingLanguage ModellingLarge Language Model+1

Variation in Coreference Strategies across Genres and Production Media

2020-12-01 · COLING 2020 8 · Berfin Akta{\c{s}}, Manfred Stede

In response to (i) inconclusive results in the literature as to the properties of coreference chains in written versus spoken language, and (ii) a general lack of work on automatic coreference resolution on both spoken l…

coreference-resolutionCoreference Resolution

Hidden Resources ― Strategies to Acquire and Exploit Potential Spoken Language Resources in National Archives

2016-05-01 · LREC 2016 5 · Jens Edlund, Joakim Gustafson

In 2014, the Swedish government tasked a Swedish agency, The Swedish Post and Telecom Authority (PTS), with investigating how to best create and populate an infrastructure for spoken language resources (Ref N2014/2840/IT…

Position

Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models

2025-05-23 · Chi-Yuan Hsiao, Ke-Han Lu, Kai-Wei Chang, Chih-Kai Yang 외

End-to-end training of Spoken Language Models (SLMs) commonly involves adapting pre-trained text-based Large Language Models (LLMs) to the speech modality through multi-stage training on diverse tasks such as ASR, TTS an…

Continual LearningQuestion Answering