paper-with-me

Speech-to-Speech Translation

3개 벤치마크 · 논문 141편 · 이 태스크의 논문 보기 →

Benchmarks

TAT

결과 8개

FLEURS X-eng

결과 7개

CVSS

결과 3개

Most implemented

Papers

Is Prosody Lost in Translation? Fine-Grained Cross-Lingual Prosody Similarity Across Languages

2026-08-28 · Haopeng Xie, Ismail Rasim Ulgen, Sofia Son, Berrak Sisman 외 arxiv

Prosody plays an important role in speech translation, conveying information such as emphasis, emotion, and intent beyond lexical content. However, despite recent progress in expressive speech-to-speech translation (S2ST…

Speech-to-Speech Translation

STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity

2026-06-24 · Sitong Cheng, Weizhen Bian, Songjun Cao, Jin Li 외 arxiv

Speech-to-speech translation (S2ST) should preserve not only lexical meaning, but also expressive attributes: emotion, scenario style (e.g., news reporting vs. dramatic dialogue), and nonverbal vocalizations (NVs). Moreo…

Speech-to-Speech Translation

A Practical Evaluation Method for Long-Form Simultaneous Speech-to-Speech Translation

2026-06-13 · Yulin Xue, Siqi Ouyang, Lei Li arxiv

Simultaneous speech-to-speech translation (SimulS2ST) enables real-time cross-lingual communication, but existing evaluation has focused largely on short or pre-segmented speech rather than long-form, continuous input. P…

Speech-to-Speech TranslationSpeech Recognition

Evaluating and Preserving Lexical Stress in English-to-Chinese Speech-to-Speech Translation

2026-06-13 · Yuchen Song, Xi Chen, Mingze Li, Satoshi Nakamura arxiv

Speech-to-speech translation (S2ST) systems have achieved impressive progress in semantic accuracy and speech naturalness. However, the cross-lingual transfer of lexical stress, a vital cue for emphasis and speaker inten…

Speech-to-Speech TranslationCross-Lingual Transfer

NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation

2026-06-11 · Dongwook Lee, Youngho Cho, Sangkwon Park, Heeseung Kim 외 arxiv

Simultaneous speech-to-speech translation aims to enable near-real-time communication by minimizing latency, offering a compelling, real-time alternative to the high latency of consecutive translation. However, the exces…

Speech-to-Speech Translation

Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data

2026-06-11 · Qixu Chen, Satoshi Nakamura arxiv

Large-scale mined corpora provide abundant training data for end-to-end speech-to-speech translation (S2ST) but may contain noise, misalignment, and semantic errors. Filtering noisy data is crucial to maintain robust spe…

Speech-to-Speech Translation

전체 141편 보기 →