paper-with-me

Papers automatic-speech-translation

“automatic-speech-translation” 태그가 달린 논문 23편 · 필터 해제

LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models

2025-06-05 · Wen Ding, Fan Qian

We introduce LESS (Large Language Model Enhanced Semi-supervised Learning), a versatile framework that leverages Large Language Models (LLMs) to correct pseudo labels generated from in-the-wild data. Within the LESS fram…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationLanguage Modeling+5

Word Level Timestamp Generation for Automatic Speech Recognition and Translation

2025-05-21 · Ke Hu, Krishna Puvvada, Elena Rastorgueva, Zhehuai Chen 외

We introduce a data-driven approach for enabling word-level timestamp prediction in the Canary model. Accurate timestamp information is crucial for a variety of downstream tasks such as speech content retrieval and timed…

Automatic Speech Recognitionautomatic-speech-translationPredictionspeech-recognition+2

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

2025-05-13 · George Saon, Avihu Dekel, Alexander Brooks, Tohru Nagano 외

Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR and automatic speech translation (AST). The models were trained by modality aligning the 2B and 8B parameter vari…

automatic-speech-translationBenchmarking

BhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages

2024-11-07 · Sparsh Jain, Ashwin Sankar, Devilal Choudhary, Dhairya Suman 외

Automatic Speech Translation (AST) datasets for Indian languages remain critically scarce, with public resources covering fewer than 10 of the 22 official languages. This scarcity has resulted in AST systems for Indian l…

automatic-speech-translationSynthetic Data GenerationTranslation

EMMeTT: Efficient Multimodal Machine Translation Training

2024-09-20 · Piotr Żelasko, Zhehuai Chen, Mengru Wang, Daniel Galvez 외

A rising interest in the modality extension of foundation language models warrants discussion on the most effective, and efficient, multimodal training approach. This work focuses on neural machine translation (NMT) and …

automatic-speech-translationDecoderMachine TranslationMultimodal Machine Translation+2

Chain-of-Thought Prompting for Speech Translation

2024-09-17 · Ke Hu, Zhehuai Chen, Chao-Han Huck Yang, Piotr Żelasko 외

Large language models (LLMs) have demonstrated remarkable advancements in language understanding and generation. Building on the success of text-based LLMs, recent research has adapted these models to use speech embeddin…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationDecoder+3

Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text

2024-09-17 · Hongfei Xue, Wei Ren, Xuelong Geng, Kun Wei 외

Integrating audio encoders with LLMs through connectors has enabled these models to process and comprehend audio modalities, significantly enhancing speech-to-text tasks, including automatic speech recognition (ASR) and …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationspeech-recognition+2

MooER: LLM-based Speech Recognition and Translation Models from Moore Threads

2024-08-09 · Junhao Xu, Zhenlin Liang, Yi Liu, Yichao Hu 외

In this paper, we present MooER, a LLM-based large-scale automatic speech recognition (ASR) / automatic speech translation (AST) model of Moore Threads. A 5000h pseudo labeled dataset containing open source and self coll…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationspeech-recognition+1

Seamless: Multilingual Expressive and Streaming Speech Translation

2023-12-08 · Seamless Communication, Loïc Barrault, Yu-An Chung, Mariano Coria Meglioli 외

Large-scale automatic speech translation systems today lack key features that help machine-mediated communication feel seamless when compared to human-to-human dialogue. In this work, we introduce a family of models that…

automatic-speech-translationMachine TranslationMultimodal Machine TranslationRed Teaming+1

Improving End-to-End Speech Translation by Imitation-Based Knowledge Distillation with Synthetic Transcripts

2023-07-17 · Rebekka Hubert, Artem Sokolov, Stefan Riezler

End-to-end automatic speech translation (AST) relies on data that combines audio inputs with text translation outputs. Previous work used existing large parallel corpora of transcriptions and translations in a knowledge …

automatic-speech-translationImitation LearningKnowledge DistillationMachine Translation+3

Improved Cross-Lingual Transfer Learning For Automatic Speech Translation

2023-06-01 · Sameer Khurana, Nauman Dawalatabad, Antoine Laurent, Luis Vicente 외

Research in multilingual speech-to-text translation is topical. Having a single model that supports multiple translation tasks is desirable. The goal of this work it to improve cross-lingual transfer learning in multilin…

automatic-speech-translationCross-Lingual TransferDecoderKnowledge Distillation+5

Robustness of Multi-Source MT to Transcription Errors

2023-05-26 · Dominik Macháček, Peter Polák, Ondřej Bojar, Raj Dabre

Automatic speech translation is sensitive to speech recognition errors, but in a multilingual scenario, the same content may be available in various languages via simultaneous interpreting, dubbing or subtitling. In this…

automatic-speech-translationMachine Translationspeech-recognitionSpeech Recognition+1

Mu$^{2}$SLAM: Multitask, Multilingual Speech and Language Models

2022-12-19 · Yong Cheng, Yu Zhang, Melvin Johnson, Wolfgang Macherey 외

We present Mu$^{2}$SLAM, a multilingual sequence-to-sequence model pre-trained jointly on unlabeled speech, unlabeled text and supervised data spanning Automatic Speech Recognition (ASR), Automatic Speech Translation (AS…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationDecoder+8

Development of Hybrid ASR Systems for Low Resource Medical Domain Conversational Telephone Speech

2022-10-24 · Christoph Lüscher, Mohammad Zeineldeen, Zijian Yang, Tina Raissi 외

Language barriers present a great challenge in our increasingly connected and global world. Especially within the medical domain, e.g. hospital or emergency room, communication difficulties and delays may lead to malprac…

automatic-speech-translationTranslation

LiSTra Automatic Speech Translation: English to Lingala Case Study

2022-06-01 · DCLRL (LREC) 2022 6 · Salomon Kabongo Kabenamualu, Vukosi Marivate, Herman Kamper

In recent years there has been great interest in addressing the data scarcity of African languages and providing baseline models for different Natural Language Processing tasks (Orife et al., 2020). Several initiatives (…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationDecoder+4

LiSTra, Automatic Speech Translation: English to Lingala case study

2021-05-16 · ACL ARR May 2021 5 · Anonymous

In recent years there have been great interests in addressing the low resourcefulness of African languages and provide baseline models for different Natural Language Processing tasks. Several initiatives on the continent…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationDecoder+4

ELITR Multilingual Live Subtitling: Demo and Strategy

2021-04-01 · EACL 2021 2 · Ond{\v{r}}ej Bojar, Dominik Mach{\'a}{\v{c}}ek, Sangeet Sagar, Otakar Smr{\v{z}} 외

This paper presents an automatic speech translation system aimed at live subtitling of conference presentations. We describe the overall architecture and key processing components. More importantly, we explain our strate…

automatic-speech-translationTranslation

Towards the evaluation of automatic simultaneous speech translation from a communicative perspective

2021-03-15 · ACL (IWSLT) 2021 8 · claudio Fantinuoli, Bianca Prandi

In recent years, automatic speech-to-speech and speech-to-text translation has gained momentum thanks to advances in artificial intelligence, especially in the domains of speech recognition and machine translation. The q…

automatic-speech-translationInformativenessMachine Translationspeech-recognition+4

Breeding Gender-aware Direct Speech Translation Systems

2020-12-09 · COLING 2020 8 · Marco Gaido, Beatrice Savoldi, Luisa Bentivogli, Matteo Negri 외

In automatic speech translation (ST), traditional cascade approaches involving separate transcription and translation steps are giving ground to increasingly competitive and more robust direct solutions. In particular, b…

automatic-speech-translationMachine TranslationTranslation

SkinAugment: Auto-Encoding Speaker Conversions for Automatic Speech Translation

2020-02-27 · Arya D. McCarthy, Liezl Puzon, Juan Pino

We propose autoencoding speaker conversion for training data augmentation in automatic speech translation. This technique directly transforms an audio sequence, resulting in audio synthesized to resemble another speaker'…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationData Augmentation+4
1–20 / 23 다음 →