paper-with-me

Papers

Leveraging Sub Label Dependencies in Code Mixed Indian Languages for Part-Of-Speech Tagging using Conditional Random Fields.

2022-06-01 · WILDRE (LREC) 2022 6 · Akash Kumar Gautam

Code-mixed text sequences often lead to challenges in the task of correct identification of Part-Of-Speech tags. However, lexical dependencies created while alternating between multiple languages can be leveraged to improve the performance of such tasks. Indian languages with rich morphological structure and highly inflected nature provide such an opportunity. In this work, we exploit these sub-label dependencies using conditional random fields (CRFs) by defining feature extraction functions on three distinct language pairs (Hindi-English, Bengali-English, and Telugu-English). Our results demonstrate a significant increase in the tagging performance if the feature extraction functions employ the rich inner structure of such languages.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Part-Of-Speech Tagging

Similar Papers 제목 키워드 기반

A CRF Based POS Tagger for Code-mixed Indian Social Media Text

2016-12-23 · Kamal Sarkar

In this work, we describe a conditional random fields (CRF) based system for Part-Of- Speech (POS) tagging of code-mixed Indian social media text as part of our participation in the tool contest on POS tagging for codemi…

Part-Of-Speech TaggingPOSPOS Tagging

Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR

2026-05-28 · Debajyoti Mazumder, Divyansh Pathak, Prashant Kodali, Aditya Joshi 외 arxiv

Large language models recall knowledge reliably in English but often fail on the same query posed in a lower-resourced language -- a crosslingual consistency gap that remains underexplored for Indian languages and their …

Speaker Verification

Sentiment Classification of Code-Mixed Tweets using Bi-Directional RNN and Language Tags

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Sainik Mahata, Dipankar Das, Sivaji Bandyopadhyay

Sentiment analysis tools and models have been developed extensively throughout the years, for European languages. In contrast, similar tools for Indian Languages are scarce. This is because, state-of-the-art pre-processi…

POSSentiment AnalysisSentiment Classification

Experiments with POS Tagging Code-mixed Indian Social Media Text

2016-10-31 · Prakash B. Pimpale, Raj Nath Patel

This paper presents Centre for Development of Advanced Computing Mumbai's (CDACM) submission to the NLP Tools Contest on Part-Of-Speech (POS) Tagging For Code-mixed Indian Social Media Text (POSCMISMT) 2015 (collocated w…

Part-Of-Speech TaggingPOSPOS TaggingTAG

Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages

2026-03-01 · Kaushal Santosh Bhogale, Tahir Javed, Greeshma Susan John, Dhruv Rathi 외 arxiv

Evaluating ASR systems for Indian languages is challenging due to spelling variations, suffix splitting flexibility, and non-standard spellings in code-mixed words. Traditional Word Error Rate (WER) often presents a blea…

Speech Recognition