paper-with-me

홈 › Papers

Prabhupadavani: A Code-mixed Speech Translation Data for 25 Languages

2022-01-27 · LaTeCHCLfL (COLING) 2022 10 · Jivnesh Sandhan, Ayush Daksh, Om Adideva Paranjay, Laxmidhar Behera, Pawan Goyal

Nowadays, the interest in code-mixing has become ubiquitous in Natural Language Processing (NLP); however, not much attention has been given to address this phenomenon for Speech Translation (ST) task. This can be solely attributed to the lack of code-mixed ST task labelled data. Thus, we introduce Prabhupadavani, which is a multilingual code-mixed ST dataset for 25 languages. It is multi-domain, covers ten language families, containing 94 hours of speech by 130+ speakers, manually aligned with corresponding text in the target language. The Prabhupadavani is about Vedic culture and heritage from Indic literature, where code-switching in the case of quotation from literature is important in the context of humanities teaching. To the best of our knowledge, Prabhupadvani is the first multi-lingual code-mixed ST dataset available in the ST literature. This data also can be used for a code-mixed machine translation task. All the dataset can be accessed at https://github.com/frozentoad9/CMST.

📄 PDF Abstract BibTeX arXiv:2201.11391

Code (1)

frozentoad9/CMST 공식 구현

Tasks

Cultural Vocal Bursts Intensity PredictionMachine TranslationTranslation

Similar Papers 제목 키워드 기반

AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation

2025-03-18 · Findings (ACL) 2021 8 · Wuwei Huang, Dexin Wang, Deyi Xiong

In end-to-end speech translation, acoustic representations learned by the encoder are usually fixed and static, from the perspective of the decoder, which is not desirable for dealing with the cross-modal and cross-lingu…

DecoderSpeech-to-TextSpeech-to-Text TranslationTranslation

"Is Hate Lost in Translation?": Evaluation of Multilingual LGBTQIA+ Hate Speech Detection

2024-10-15 · Fai Leui Chan, Duke Nguyen, Aditya Joshi

This paper explores the challenges of detecting LGBTQIA+ hate speech of large language models across multiple languages, including English, Italian, Chinese and (code-switched) English-Tamil, examining the impact of mach…

Hate Speech DetectionMachine TranslationTranslation

Mixed-Precision Training for NLP and Speech Recognition with OpenSeq2Seq

2018-05-25 · Oleksii Kuchaiev, Boris Ginsburg, Igor Gitman, Vitaly Lavrukhin 외

We present OpenSeq2Seq - a TensorFlow-based toolkit for training sequence-to-sequence models that features distributed and mixed-precision training. Benchmarks on machine translation and speech recognition tasks show tha…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+3

Domain Curricula for Code-Switched MT at MixMT 2022

2022-10-31 · Lekan Raheem, Maab Elrashid

In multilingual colloquial settings, it is a habitual occurrence to compose expressions of text or speech containing tokens or phrases of different languages, a phenomenon popularly known as code-switching or code-mixing…

Machine TranslationSentenceTranslation

Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing

2024-06-04 · Viet Anh Trinh, Rosy Southwell, Yiwen Guan, Xinlu He 외

Recent work on discrete speech tokenization has paved the way for models that can seamlessly perform multiple tasks across modalities, e.g., speech recognition, text to speech, speech to speech translation. Moreover, lar…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+6