paper-with-me

홈 › Papers

M2H-GAN: A GAN-based Mapping from Machine to Human Transcripts for Speech Understanding

2019-04-13 · Titouan Parcollet, Mohamed Morchid, Xavier Bost, Georges Linarès

Deep learning is at the core of recent spoken language understanding (SLU) related tasks. More precisely, deep neural networks (DNNs) drastically increased the performances of SLU systems, and numerous architectures have been proposed. In the real-life context of theme identification of telephone conversations, it is common to hold both a human, manual (TRS) and an automatically transcribed (ASR) versions of the conversations. Nonetheless, and due to production constraints, only the ASR transcripts are considered to build automatic classifiers. TRS transcripts are only used to measure the performances of ASR systems. Moreover, the recent performances in term of classification accuracy, obtained by DNN related systems are close to the performances reached by humans, and it becomes difficult to further increase the performances by only considering the ASR transcripts. This paper proposes to distillates the TRS knowledge available during the training phase within the ASR representation, by using a new generative adversarial network called M2H-GAN to generate a TRS-like version of an ASR document, to improve the theme identification performances.

📄 PDF Abstract BibTeX arXiv:1905.01957

Code (0)

등록된 구현이 없습니다.

Tasks

Generative Adversarial NetworkSpoken Language Understanding

Similar Papers 제목 키워드 기반

Back Translation for Speech-to-text Translation Without Transcripts

2023-05-15 · Qingkai Fang, Yang Feng

The success of end-to-end speech-to-text translation (ST) is often achieved by utilizing source transcripts, e.g., by pre-training with automatic speech recognition (ASR) and machine translation (MT) tasks, or by introdu…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)de-enMachine Translation+5

Do We Still Need Automatic Speech Recognition for Spoken Language Understanding?

2021-11-29 · Lasse Borgholt, Jakob Drachmann Havtorn, Mostafa Abdou, Joakim Edin 외

Spoken language understanding (SLU) tasks are usually solved by first transcribing an utterance with automatic speech recognition (ASR) and then feeding the output to a text-based model. Recent advances in self-supervise…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationnamed-entity-recognition+7

Quantification of stylistic differences in human- and ASR-produced transcripts of African American English

2024-09-04 · Annika Heuser, Tyler Kendall, Miguel Del Rio, Quinten McNamara 외

Common measures of accuracy used to assess the performance of automatic speech recognition (ASR) systems, as well as human transcribers, conflate multiple sources of error. Stylistic differences, such as verbatim vs non-…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Robust Translation of French Live Speech Transcripts

2022-09-01 · AMTA 2022 9 · Elise Bertin-Lemée, Guillaume Klein, Josep Crego, Jean Senellart

Despite a narrowed performance gap with direct approaches, cascade solutions, involving automatic speech recognition (ASR) and machine translation (MT) are still largely employed in speech translation (ST). Direct approa…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSentence+3

Robust Neural Machine Translation for Clean and Noisy Speech Transcripts

2019-10-22 · EMNLP (IWSLT) 2019 11 · Mattia Antonino Di Gangi, Robert Enyedi, Alessandra Brusadin, Marcello Federico

Neural machine translation models have shown to achieve high quality when trained and fed with well structured and punctuated input texts. Unfortunately, the latter condition is not met in spoken language translation, wh…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationNMT+3