paper-with-me

홈 › Papers

Who Are We Talking About? Handling Person Names in Speech Translation

2022-05-13 · IWSLT (ACL) 2022 5 · Marco Gaido, Matteo Negri, Marco Turchi

Recent work has shown that systems for speech translation (ST) -- similarly to automatic speech recognition (ASR) -- poorly handle person names. This shortcoming does not only lead to errors that can seriously distort the meaning of the input, but also hinders the adoption of such systems in application scenarios (like computer-assisted interpreting) where the translation of named entities, like person names, is crucial. In this paper, we first analyse the outputs of ASR/ST systems to identify the reasons of failures in person name transcription/translation. Besides the frequency in the training data, we pinpoint the nationality of the referred person as a key factor. We then mitigate the problem by creating multilingual models, and further improve our ST systems by forcing them to jointly generate transcripts and translations, prioritising the former over the latter. Overall, our solutions result in a relative improvement in token-level person name accuracy by 47.8% on average for three language pairs (en->es,fr,it).

📄 PDF Abstract BibTeX arXiv:2205.06755

Code (1)

hlt-mt/fbk-fairseq 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech RecognitionTranslation

Similar Papers 제목 키워드 기반

Who Are We Talking About? Handling Person Names in Speech Translation

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Recent work has shown that systems for speech translation (ST) -- similarly to automatic speech recognition (ASR) -- poorly handle person names. This shortcoming does not only lead to errors that can seriously distort th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person

2021-08-09 · Xinsheng Wang, Qicong Xie, Jihua Zhu, Lei Xie 외

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. In this paper, we present an automatic me…

Talking Head Generationtext-to-speechText to Speech

AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial Animation

2023-10-11 · Liyang Chen, Weihong Bao, Shun Lei, Boshi Tang 외

Speech-driven 3D facial animation aims at generating facial movements that are synchronized with the driving speech, which has been widely explored recently. Existing works mostly neglect the person-specific talking styl…

An Alternative to Low-level-Sychrony-Based Methods for Speech Detection

2010-12-01 · NeurIPS 2010 12 · Javier R. Movellan, Paul L. Ruvolo

Determining whether someone is talking has applications in many areas such as speech recognition, speaker diarization, social robotics, facial expression recognition, and human computer interaction. One popular approach…

Facial Expression RecognitionFacial Expression Recognition (FER)speaker-diarizationSpeaker Diarization+2

Retraining-free Customized ASR for Enharmonic Words Based on a Named-Entity-Aware Model and Phoneme Similarity Estimation

2023-05-29 · Yui Sudo, Kazuya Hata, Kazuhiro Nakadai

End-to-end automatic speech recognition (E2E-ASR) has the potential to improve performance, but a specific issue that needs to be addressed is the difficulty it has in handling enharmonic words: named entities (NEs) with…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition