Papers automatic-speech-translation
“automatic-speech-translation” 태그가 달린 논문 23편 · 필터 해제
LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models
We introduce LESS (Large Language Model Enhanced Semi-supervised Learning), a versatile framework that leverages Large Language Models (LLMs) to correct pseudo labels generated from in-the-wild data. Within the LESS fram…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationLanguage Modeling+5Word Level Timestamp Generation for Automatic Speech Recognition and Translation
We introduce a data-driven approach for enabling word-level timestamp prediction in the Canary model. Accurate timestamp information is crucial for a variety of downstream tasks such as speech content retrieval and timed…
Automatic Speech Recognitionautomatic-speech-translationPredictionspeech-recognition+2Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR and automatic speech translation (AST). The models were trained by modality aligning the 2B and 8B parameter vari…
automatic-speech-translationBenchmarkingBhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages
Automatic Speech Translation (AST) datasets for Indian languages remain critically scarce, with public resources covering fewer than 10 of the 22 official languages. This scarcity has resulted in AST systems for Indian l…
automatic-speech-translationSynthetic Data GenerationTranslationEMMeTT: Efficient Multimodal Machine Translation Training
A rising interest in the modality extension of foundation language models warrants discussion on the most effective, and efficient, multimodal training approach. This work focuses on neural machine translation (NMT) and …
automatic-speech-translationDecoderMachine TranslationMultimodal Machine Translation+2Chain-of-Thought Prompting for Speech Translation
Large language models (LLMs) have demonstrated remarkable advancements in language understanding and generation. Building on the success of text-based LLMs, recent research has adapted these models to use speech embeddin…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationDecoder+3Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text
Integrating audio encoders with LLMs through connectors has enabled these models to process and comprehend audio modalities, significantly enhancing speech-to-text tasks, including automatic speech recognition (ASR) and …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationspeech-recognition+2MooER: LLM-based Speech Recognition and Translation Models from Moore Threads
In this paper, we present MooER, a LLM-based large-scale automatic speech recognition (ASR) / automatic speech translation (AST) model of Moore Threads. A 5000h pseudo labeled dataset containing open source and self coll…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationspeech-recognition+1Seamless: Multilingual Expressive and Streaming Speech Translation
Large-scale automatic speech translation systems today lack key features that help machine-mediated communication feel seamless when compared to human-to-human dialogue. In this work, we introduce a family of models that…
automatic-speech-translationMachine TranslationMultimodal Machine TranslationRed Teaming+1Improving End-to-End Speech Translation by Imitation-Based Knowledge Distillation with Synthetic Transcripts
End-to-end automatic speech translation (AST) relies on data that combines audio inputs with text translation outputs. Previous work used existing large parallel corpora of transcriptions and translations in a knowledge …
automatic-speech-translationImitation LearningKnowledge DistillationMachine Translation+3Improved Cross-Lingual Transfer Learning For Automatic Speech Translation
Research in multilingual speech-to-text translation is topical. Having a single model that supports multiple translation tasks is desirable. The goal of this work it to improve cross-lingual transfer learning in multilin…
automatic-speech-translationCross-Lingual TransferDecoderKnowledge Distillation+5Robustness of Multi-Source MT to Transcription Errors
Automatic speech translation is sensitive to speech recognition errors, but in a multilingual scenario, the same content may be available in various languages via simultaneous interpreting, dubbing or subtitling. In this…
automatic-speech-translationMachine Translationspeech-recognitionSpeech Recognition+1Mu$^{2}$SLAM: Multitask, Multilingual Speech and Language Models
We present Mu$^{2}$SLAM, a multilingual sequence-to-sequence model pre-trained jointly on unlabeled speech, unlabeled text and supervised data spanning Automatic Speech Recognition (ASR), Automatic Speech Translation (AS…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationDecoder+8Development of Hybrid ASR Systems for Low Resource Medical Domain Conversational Telephone Speech
Language barriers present a great challenge in our increasingly connected and global world. Especially within the medical domain, e.g. hospital or emergency room, communication difficulties and delays may lead to malprac…
automatic-speech-translationTranslationLiSTra Automatic Speech Translation: English to Lingala Case Study
In recent years there has been great interest in addressing the data scarcity of African languages and providing baseline models for different Natural Language Processing tasks (Orife et al., 2020). Several initiatives (…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationDecoder+4LiSTra, Automatic Speech Translation: English to Lingala case study
In recent years there have been great interests in addressing the low resourcefulness of African languages and provide baseline models for different Natural Language Processing tasks. Several initiatives on the continent…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationDecoder+4ELITR Multilingual Live Subtitling: Demo and Strategy
This paper presents an automatic speech translation system aimed at live subtitling of conference presentations. We describe the overall architecture and key processing components. More importantly, we explain our strate…
automatic-speech-translationTranslationTowards the evaluation of automatic simultaneous speech translation from a communicative perspective
In recent years, automatic speech-to-speech and speech-to-text translation has gained momentum thanks to advances in artificial intelligence, especially in the domains of speech recognition and machine translation. The q…
automatic-speech-translationInformativenessMachine Translationspeech-recognition+4Breeding Gender-aware Direct Speech Translation Systems
In automatic speech translation (ST), traditional cascade approaches involving separate transcription and translation steps are giving ground to increasingly competitive and more robust direct solutions. In particular, b…
automatic-speech-translationMachine TranslationTranslationSkinAugment: Auto-Encoding Speaker Conversions for Automatic Speech Translation
We propose autoencoding speaker conversion for training data augmentation in automatic speech translation. This technique directly transforms an audio sequence, resulting in audio synthesized to resemble another speaker'…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationData Augmentation+4