Speech-to-Text Translation
11개 벤치마크 · 논문 176편 · 이 태스크의 논문 보기 →
Benchmarks
MuST-C EN->DE
MuST-C EN->ES
MuST-C EN->FR
CoVoST 2 X-eng
CoVoST 2 eng-X
FLEURS X-eng
FLEURS eng-X
MediBeng
MuST-C
libri-trans
MuST-C EN->NL
Most implemented
fairseq S2T: Fast Speech-to-Text Modeling with fairseq
SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit
A$^3$T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and Editing
SHAS: Approaching optimal Segmentation for End-to-End Speech Translation
Papers
NOPE-HYPE: A Structured Simulation Workflow for Robust Speech-to-Text Across Diverse Acoustic Environments
Robust speech-to-text translation systems should perform reliably across diverse acoustic conditions, yet practical pipelines lack controllable tools for systematic environment exploration. Large speech models remain sen…
Speech-to-Text TranslationSPRI: SVD-Partitioned Residual Initialization for Data-Constrained MoE Upcycling
Mixture-of-Experts (MoE) models enable efficient scaling, but training them from scratch remains prohibitively expensive. MoE upcycling mitigates this cost by converting pretrained dense models into sparse MoE models. Ho…
Speech-to-Text TranslationA Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026
We implement simultaneous translation capability with the offline direct speech-to-text translation model Canary, using the state-of-the-art policy AlignAtt, and submit it to IWSLT 2026 Simultaneous Speech Translation Sh…
Speech-to-Text TranslationOpenSTBench: Beyond Semantic Evaluation for Speech Translation
Speech translation systems increasingly span speech-to-text translation (S2TT), speech-to-speech translation (S2ST), offline translation, and streaming generation, producing outputs that differ in modality, speech realiz…
Speech-to-Speech TranslationSpeech-to-Text TranslationDOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
Simultaneous speech-to-text translation (SimulST) generates translations while speech is still unfolding, requiring a streaming policy that decides when to read and when to write. State-of-the-art approaches rely on atte…
Speech-to-Text TranslationParameter- and Bandwidth-Efficient Edge--cloud Many-to-Many Speech-to-Text Translation
Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment paradigms face critical challenges: pure on-device models suffer from r…
Speech-to-Text TranslationSpeech Recognition