paper-with-me

홈 › Papers

ILID: Native Script Language Identification for Indian Languages

2025-07-16 · Yash Ingle, Pruthwik Mishra arxiv

The language identification task is a crucial fundamental step in NLP. Often it serves as a pre-processing step for widely used NLP applications such as multilingual machine translation, information retrieval, question and answering, and text summarization. The core challenge of language identification lies in distinguishing languages in noisy, short, and code-mixed environments. This becomes even harder in case of diverse Indian languages that exhibit lexical and phonetic similarities, but have distinct differences. Many Indian languages share the same script, making the task even more challenging. Taking all these challenges into account, we develop and release a dataset of 250K sentences consisting of 23 languages including English and all 22 official Indian languages labeled with their language identifiers, where data in most languages are newly created. We also develop and release baseline models using state-of-the-art approaches in machine learning and fine-tuning pre-trained transformer models. Our models outperforms the state-of-the-art pre-trained transformer models for the language identification task. The dataset and the codes are available at https://yashingle-ai.github.io/ILID/ and in Huggingface open source libraries.

📄 PDF Abstract BibTeX arXiv:2507.11832

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationInformation RetrievalMachine TranslationText Summarization

Similar Papers 제목 키워드 기반

Bhasha-Abhijnaanam: Native-script and romanized Language Identification for 22 Indic languages

2023-05-25 · Yash Madhani, Mitesh M. Khapra, Anoop Kunchukuttan

We create publicly available language identification (LID) datasets and models in all 22 Indian languages listed in the Indian constitution in both native-script and romanized text. First, we create Bhasha-Abhijnaanam, a…

Language Identification

Annotation Efficient Language Identification from Weak Labels

2020-11-01 · EMNLP (WNUT) 2020 11 · Shriphani Palakodety, Ashiqur KhudaBukhsh

India is home to several languages with more than 30m speakers. These languages exhibit significant presence on social media platforms. However, several of these widely-used languages are under-addressed by current Natur…

Language Identification

DuDe: Dual-Decoder Multilingual ASR for Indian Languages using Common Label Set

2022-10-30 · Arunkumar A, Mudit Batra, Umesh S

In a multilingual country like India, multilingual Automatic Speech Recognition (ASR) systems have much scope. Multilingual ASR systems exhibit many advantages like scalability, maintainability, and improved performance …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognition+2

An Overview of Indian Spoken Language Recognition from Machine Learning Perspective

2022-11-30 · Spandan Dey, Md Sahidullah, Goutam Saha

Automatic spoken language identification (LID) is a very important research field in the era of multilingual voice-command-based human-computer interaction (HCI). A front-end LID module helps to improve the performance o…

Language IdentificationSpoken language identification

Identification of Indian Languages using Ghost-VLAD pooling

2020-02-05 · Krishna D N, Ankita Patil, M. S. P Raj, Sai Prasad H S 외

In this work, we propose a new pooling strategy for language identification by considering Indian languages. The idea is to obtain utterance level features for any variable length audio for robust language recognition. W…

Language Identification