paper-with-me

Papers

CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving

2024-06-16 · Bhavani Shankar, Preethi Jyothi, Pushpak Bhattacharyya

Code-switching is a widely prevalent linguistic phenomenon in multilingual societies like India. Building speech-to-text models for code-switched speech is challenging due to limited availability of datasets. In this work, we focus on the problem of spoken translation (ST) of code-switched speech in Indian languages to English text. We present a new end-to-end model architecture COSTA that scaffolds on pretrained automatic speech recognition (ASR) and machine translation (MT) modules (that are more widely available for many languages). Speech and ASR text representations are fused using an aligned interleaving scheme and are fed further as input to a pretrained MT module; the whole pipeline is then trained end-to-end for spoken translation using synthetically created ST data. We also release a new evaluation benchmark for code-switched Bengali-English, Hindi-English, Marathi-English and Telugu- English speech to English text. COSTA significantly outperforms many competitive cascaded and end-to-end multimodal baselines by up to 3.5 BLEU points.

📄 PDF Abstract BibTeX arXiv:2406.10993

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognitionSpeech RecognitionSpeech-to-TextTranslation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

ArzEn-ST: A Three-way Speech Translation Corpus for Code-Switched Egyptian Arabic - English

2022-11-22 · Injy Hamed, Nizar Habash, Slim Abdennadher, Ngoc Thang Vu

We present our work on collecting ArzEn-ST, a code-switched Egyptian Arabic - English Speech Translation Corpus. This corpus is an extension of the ArzEn speech corpus, which was collected through informal interviews wit…

Machine TranslationTranslation

CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset

2025-09-17 · Brian Yan, Injy Hamed, Shuichiro Shimizu, Vasista Lodagala 외 arxiv

We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. CS-FLEURS consists of 4 test sets which cover in total 113 uniqu…

Speech Recognition

End-to-End Code Switching Language Models for Automatic Speech Recognition

2020-06-16 · Ahan M. R., Shreyas Sunil Kulkarni

In this paper, we particularly work on the code-switched text, one of the most common occurrences in the bilingual communities across the world. Due to the discrepancies in the extraction of code-switched text from an Au…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+2

CoVoSwitch: Machine Translation of Synthetic Code-Switched Text Based on Intonation Units

2024-07-19 · Yeeun Kang

Multilingual code-switching research is often hindered by the lack and linguistically biased status of available datasets. To expand language representation, we synthesize code-switching data by replacing intonation unit…

Machine TranslationSpeech-to-TextSpeech-to-Text TranslationTranslation

ArzEn-LLM: Code-Switched Egyptian Arabic-English Translation and Speech Recognition Using LLMs

2024-06-26 · Ahmed Heakl, Youssef Zaghloul, Mennatullah Ali, Rania Hossam 외

Motivated by the widespread increase in the phenomenon of code-switching between Egyptian Arabic and English in recent times, this paper explores the intricacies of machine translation (MT) and automatic speech recogniti…

ArzEn Code-switched Translation to araArzEn Code-switched Translation to engArzEn Speech RecognitionAutomatic Speech Recognition+9