BLASER: A Text-Free Speech-to-Speech Translation Evaluation Metric
End-to-End speech-to-speech translation (S2ST) is generally evaluated with text-based metrics. This means that generated speech has to be automatically transcribed, making the evaluation dependent on the availability and quality of automatic speech recognition (ASR) systems. In this paper, we propose a text-free evaluation metric for end-to-end S2ST, named BLASER, to avoid the dependency on ASR systems. BLASER leverages a multilingual multimodal encoder to directly encode the speech segments for source input, translation output and reference into a shared embedding space and computes a score of the translation quality that can be used as a proxy to human evaluation. To evaluate our approach, we construct training and evaluation sets from more than 40k human annotations covering seven language directions. The best results of BLASER are achieved by training with supervision from human rating scores. We show that when evaluated at the sentence level, BLASER correlates significantly better with human judgment compared to ASR-dependent metrics including ASR-SENTBLEU in all translation directions and ASR-COMET in five of them. Our analysis shows combining speech and text as inputs to BLASER does not increase the correlation with human scores, but best correlations are achieved when using speech, which motivates the goal of our research. Moreover, we show that using ASR for references is detrimental for text-based metrics.
Code (1)
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)SentenceSpeech RecognitionSpeech-to-Speech TranslationTranslationSimilar Papers 제목 키워드 기반
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
Cascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumulate. In this paper, we propose a transd…
Language ModelingLanguage ModellingMachine TranslationSpeech-to-Speech Translation+3From Flat Language Labels to Typological Priors: Structured Language Conditioning for Multilingual Speech-to-Speech Translation
Compositional speech-to-speech translation (S2ST) systems built upon speech large language models (SpeechLLMs) have recently shown promising performance. However, existing S2ST systems often either neglect source-languag…
Speech-to-Speech TranslationRepresentation Purification for End-to-End Speech Translation
Speech-to-text translation (ST) is a cross-modal task that involves converting spoken language into text in a different language. Previous research primarily focused on enhancing speech translation by facilitating knowle…
Machine TranslationRhythmSpeech-to-TextSpeech-to-Text Translation+2SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations
We present SpeechMatrix, a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. It contains speech alignments in 136 language pairs with a total of 41…
Mixture-of-ExpertsSpeech-to-Speech TranslationTranslationThe IWSLT 2018 Evaluation Campaign
The International Workshop of Spoken Language Translation (IWSLT) 2018 Evaluation Campaign featured two tasks: low-resource machine translation and speech translation. In the first task, manually transcribed speech had t…
Machine TranslationTranslation