paper-with-me

홈 › Papers

Streaming Translation and Transcription Through Speech-to-Text Causal Alignment

2026-03-12 · Roman Koshkin, Jeon Haesung, Lianbo Liu, Hao Shi, Mengjie Zhao, Yusuke Fujita, Yui Sudo arxiv

Simultaneous machine translation (SiMT) has traditionally relied on offline machine translation models coupled with human-engineered heuristics or learned policies. We propose Hikari, a policy-free, fully end-to-end model that performs simultaneous speech-to-text translation and streaming transcription by encoding READ/WRITE decisions into a probabilistic WAIT token mechanism. We also introduce Decoder Time Dilation, a mechanism that reduces autoregressive overhead and ensures a balanced training distribution. Additionally, we present a supervised fine-tuning strategy that trains the model to recover from delays, significantly improving the quality-latency trade-off. Evaluated on English-to-Japanese, German, and Russian, Hikari achieves new state-of-the-art BLEU scores in both low- and high-latency regimes, outperforming recent baselines.

📄 PDF Abstract BibTeX arXiv:2603.11578

Code (0)

등록된 구현이 없습니다.

Tasks

Speech-to-Text TranslationMachine Translation

Similar Papers 제목 키워드 기반

Turning Whisper into Real-Time Transcription System

2023-07-27 · Dominik Macháček, Raj Dabre, Ondřej Bojar

Whisper is one of the recent state-of-the-art multilingual speech recognition and translation models, however, it is not designed for real time transcription. In this paper, we build on top of Whisper and create Whisper-…

speech-recognitionSpeech RecognitionTranslation

Towards Real-World Streaming Speech Translation for Code-Switched Speech

2023-10-19 · Belen Alastruey, Matthias Sperber, Christian Gollan, Dominic Telaar 외

Code-switching (CS), i.e. mixing different languages in a single sentence, is a common phenomenon in communication and can be challenging in many Natural Language Processing (NLP) settings. Previous studies on CS speech …

SentenceTranslation

A Weakly-Supervised Streaming Multilingual Speech Model with Truly Zero-Shot Capability

2022-11-04 · Jian Xue, Peidong Wang, Jinyu Li, Eric Sun

In this paper, we introduce our work of building a Streaming Multilingual Speech Model (SM2), which can transcribe or translate multiple spoken languages into texts of the target language. The backbone of SM2 is Transfor…

Machine Translationspeech-recognitionSpeech RecognitionTranslation

Overcoming Latency Bottlenecks in On-Device Speech Translation: A Cascaded Approach with Alignment-Based Streaming MT

2025-08-18 · Zeeshan Ahmed, Frank Seide, Niko Moritz, Ju Lin 외 arxiv

This paper tackles several challenges that arise when integrating Automatic Speech Recognition (ASR) and Machine Translation (MT) for real-time, on-device streaming speech translation. Although state-of-the-art ASR syste…

Machine TranslationSpeech Recognition

Leveraging Timestamp Information for Serialized Joint Streaming Recognition and Translation

2023-10-23 · Sara Papi, Peidong Wang, Junkun Chen, Jian Xue 외

The growing need for instant spoken language transcription and translation is driven by increased global communication and cross-lingual interactions. This has made offering translations in multiple languages essential f…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderde-en+3