paper-with-me

Papers

Zero-resource Speech Translation and Recognition with LLMs

2024-12-24 · Karel Mundnich, Xing Niu, Prashant Mathur, Srikanth Ronanki, Brady Houston, Veera Raghavendra Elluru, Nilaksh Das, Zejiang Hou, Goeric Huybrechts, Anshu Bhatia, Daniel Garcia-Romero, Kyu J. Han, Katrin Kirchhoff

Despite recent advancements in speech processing, zero-resource speech translation (ST) and automatic speech recognition (ASR) remain challenging problems. In this work, we propose to leverage a multilingual Large Language Model (LLM) to perform ST and ASR in languages for which the model has never seen paired audio-text data. We achieve this by using a pre-trained multilingual speech encoder, a multilingual LLM, and a lightweight adaptation module that maps the audio representations to the token embedding space of the LLM. We perform several experiments both in ST and ASR to understand how to best train the model and what data has the most impact on performance in previously unseen languages. In ST, our best model is capable to achieve BLEU scores over 23 in CoVoST2 for two previously unseen languages, while in ASR, we achieve WERs of up to 28.2\%. We finally show that the performance of our system is bounded by the ability of the LLM to output text in the desired language.

📄 PDF Abstract BibTeX arXiv:2412.18566

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage ModellingLarge Language Modelspeech-recognitionSpeech RecognitionTranslation

Similar Papers 제목 키워드 기반

Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios

2025-05-30 · Gerard I. Gállego, Oriol Pareras, Martí Cortada Garcia, Lucas Takanori 외

We propose a Speech-to-Text Translation (S2TT) approach that integrates phoneme representations into a Chain-of-Thought (CoT) framework to improve translation in low-resource and zero-resource settings. By introducing ph…

Cross-Lingual TransferPhoneme RecognitionSpeech-to-TextSpeech-to-Text Translation+1

XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception

2024-03-21 · Hyojung Han, Mohamed Anwar, Juan Pino, Wei-Ning Hsu 외

Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. Augmenting these systems with visual signals has the potential to improve robustness to noise. Howe…

Audio-Visual Speech RecognitionRepresentation LearningRobust Speech Recognitionspeech-recognition+4

Zero-shot Cross-lingual NER via Mitigating Language Difference: An Entity-aligned Translation Perspective

2025-09-01 · Zhihao Zhang, Sophia Yat Mei Lee, Dong Zhang, Shoushan Li 외 arxiv

Cross-lingual Named Entity Recognition (CL-NER) aims to transfer knowledge from high-resource languages to low-resource languages. However, existing zero-shot CL-NER (ZCL-NER) approaches primarily focus on Latin script l…

Cross-Lingual NEREntity Alignment

Cross-Lingual Conversational Speech Summarization with Large Language Models

2024-08-12 · Max Nelson, Shannon Wotherspoon, Francis Keith, William Hartmann 외

Cross-lingual conversational speech summarization is an important problem, but suffers from a dearth of resources. While transcriptions exist for a number of languages, translated conversational speech is rare and datase…

Machine Translationspeech-recognitionSpeech RecognitionTranslation

Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models

2022-10-13 · Haoyu Wang, Wei-Qiang Zhang, Hongbin Suo, Yulong Wan

Labeled audio data is insufficient to build satisfying speech recognition systems for most of the languages in the world. There have been some zero-resource methods trying to perform phoneme or word-level speech recognit…

Language ModelingLanguage ModellingPhoneme Recognitionspeech-recognition+1