paper-with-me

홈 › Papers

SpeechT: Findings of the First Mentorship in Speech Translation

2025-02-17 · Yasmin Moslem, Juan Julián Cea Morán, Mariano Gonzalez-Gomez, Muhammad Hazim Al Farouq, Farah Abdou, Satarupa Deb

This work presents the details and findings of the first mentorship in speech translation (SpeechT), which took place in December 2024 and January 2025. To fulfil the mentorship requirements, the participants engaged in key activities, including data preparation, modelling, and advanced research. The participants explored data augmentation techniques and compared end-to-end and cascaded speech translation systems. The projects covered various languages other than English, including Arabic, Bengali, Galician, Indonesian, Japanese, and Spanish.

📄 PDF Abstract BibTeX arXiv:2502.12050

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationTranslation

Similar Papers 제목 키워드 기반

SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

2021-10-14 · ACL 2022 5 · Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang 외

Motivated by the success of T5 (Text-To-Text Transfer Transformer) in pre-trained natural language processing models, we propose a unified-modal SpeechT5 framework that explores the encoder-decoder pre-training for self-…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderQuantization+7

Speech Translation with Large Language Models: An Industrial Practice

2023-12-21 · Zhichao Huang, Rong Ye, Tom Ko, Qianqian Dong 외

Given the great success of large language models (LLMs) across various tasks, in this paper, we introduce LLM-ST, a novel and effective speech translation model constructed upon a pre-trained LLM. By integrating the larg…

Language ModelingLanguage ModellingLarge Language ModelTranslation

SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

2023-08-31 · Xin Zhang, Dong Zhang, ShiMin Li, Yaqian Zhou 외

Current speech large language models build upon discrete speech representations, which can be categorized into semantic tokens and acoustic tokens. However, existing speech tokens are not specifically designed for speech…

DecoderLanguage ModelingLanguage ModellingQuantization+2

SpeechTaxi: On Multilingual Semantic Speech Classification

2024-09-10 · Lennart Keller, Goran Glavaš

Recent advancements in multilingual speech encoding as well as transcription raise the question of the most effective approach to semantic speech classification. Concretely, can (1) end-to-end (E2E) classifiers obtained …

ClassificationCross-Lingual Transfer

PolyVoice: Language Models for Speech to Speech Translation

2023-06-05 · Qianqian Dong, Zhiying Huang, Qiao Tian, Chen Xu 외

We propose PolyVoice, a language model-based framework for speech-to-speech translation (S2ST) system. Our framework consists of two language models: a translation language model and a speech synthesis language model. We…

Language ModelingLanguage ModellingSpeech SynthesisSpeech-to-Speech Translation+1