paper-with-me

홈 › Papers

Cross-Lingual Text Classification of Transliterated Hindi and Malayalam

2021-08-31 · Jitin Krishnan, Antonios Anastasopoulos, Hemant Purohit, Huzefa Rangwala

Transliteration is very common on social media, but transliterated text is not adequately handled by modern neural models for various NLP tasks. In this work, we combine data augmentation approaches with a Teacher-Student training scheme to address this issue in a cross-lingual transfer setting for fine-tuning state-of-the-art pre-trained multilingual language models such as mBERT and XLM-R. We evaluate our method on transliterated Hindi and Malayalam, also introducing new datasets for benchmarking on real-world scenarios: one on sentiment classification in transliterated Malayalam, and another on crisis tweet classification in transliterated Hindi and Malayalam (related to the 2013 North India and 2018 Kerala floods). Our method yielded an average improvement of +5.6% on mBERT and +4.7% on XLM-R in F1 scores over their strong baselines.

📄 PDF Abstract BibTeX arXiv:2108.13620

Code (1)

jitinkrishnan/transliteration-hindi-malayalam 공식 구현 pytorch

Tasks

BenchmarkingClassificationCross-Lingual TransferData AugmentationSentiment AnalysisSentiment Classificationtext-classificationText ClassificationTransliterationXLM-R

Methods 이 논문이 사용한 방법론

XLM-R XLM-R
mBERT mBERT

Similar Papers 제목 키워드 기반

IITR-CIOL@NLU of Devanagari Script Languages 2025: Multilingual Hate Speech Detection and Target Identification in Devanagari-Scripted Languages

2024-12-23 · Siddhant Gupta, Siddh Singhal, Azmine Toushik Wasi

This work focuses on two subtasks related to hate speech detection and target identification in Devanagari-scripted languages, specifically Hindi, Marathi, Nepali, Bhojpuri, and Sanskrit. Subtask B involves detecting hat…

Binary ClassificationDiversityHate Speech Detection

Normalization and Back-Transliteration for Code-Switched Data

2021-06-01 · NAACL (CALCS) 2021 6 · Dwija Parikh, Thamar Solorio

Code-switching is an omnipresent phenomenon in multilingual communities all around the world but remains a challenge for NLP systems due to the lack of proper data and processing techniques. Hindi-English code-switched t…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Part-Of-Speech Tagging+1

Language Lexicons for Hindi-English Multilingual Text Processing

2021-06-29 · Mohd Zeeshan Ansari, Tanvir Ahmad, Noaima Bari

Language Identification in textual documents is the process of automatically detecting the language contained in a document based on its content. The present Language Identification techniques presume that a document con…

Language Identification

Leveraging Transformers for Hate Speech Detection in Conversational Code-Mixed Tweets

2021-12-18 · Zaki Mustafa Farooqi, Sreyan Ghosh, Rajiv Ratn Shah

In the current era of the internet, where social media platforms are easily accessible for everyone, people often have to deal with threats, identity attacks, hate, and bullying due to their association with a cast, cree…

Hate Speech Detection

Building English ASR model with regional language support

2025-03-10 · Purvi Agrawal, Vikas Joshi, Bharati Patidar, Ankur Gupta 외

In this paper, we present a novel approach to developing an English Automatic Speech Recognition (ASR) system that can effectively handle Hindi queries, without compromising its performance on English. We propose a novel…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3