Speech-to-Speech Translation
3개 벤치마크 · 논문 141편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Robust Speech Recognition via Large-Scale Weak Supervision
AudioLM: a Language Modeling Approach to Audio Generation
SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Textless Speech-to-Speech Translation With Limited Parallel Data
Papers
Is Prosody Lost in Translation? Fine-Grained Cross-Lingual Prosody Similarity Across Languages
Prosody plays an important role in speech translation, conveying information such as emphasis, emotion, and intent beyond lexical content. However, despite recent progress in expressive speech-to-speech translation (S2ST…
Speech-to-Speech TranslationSTEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity
Speech-to-speech translation (S2ST) should preserve not only lexical meaning, but also expressive attributes: emotion, scenario style (e.g., news reporting vs. dramatic dialogue), and nonverbal vocalizations (NVs). Moreo…
Speech-to-Speech TranslationA Practical Evaluation Method for Long-Form Simultaneous Speech-to-Speech Translation
Simultaneous speech-to-speech translation (SimulS2ST) enables real-time cross-lingual communication, but existing evaluation has focused largely on short or pre-segmented speech rather than long-form, continuous input. P…
Speech-to-Speech TranslationSpeech RecognitionEvaluating and Preserving Lexical Stress in English-to-Chinese Speech-to-Speech Translation
Speech-to-speech translation (S2ST) systems have achieved impressive progress in semantic accuracy and speech naturalness. However, the cross-lingual transfer of lexical stress, a vital cue for emphasis and speaker inten…
Speech-to-Speech TranslationCross-Lingual TransferNaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation
Simultaneous speech-to-speech translation aims to enable near-real-time communication by minimizing latency, offering a compelling, real-time alternative to the high latency of consecutive translation. However, the exces…
Speech-to-Speech TranslationLeveraging Audio-LLMs to Filter Speech-to-Speech Training Data
Large-scale mined corpora provide abundant training data for end-to-end speech-to-speech translation (S2ST) but may contain noise, misalignment, and semantic errors. Filtering noisy data is crucial to maintain robust spe…
Speech-to-Speech Translation