paper-with-me

홈 › Papers

Language Identification and Normalization of Code Mixed English and Punjabi Text

2020-12-01 · ICON 2020 12 · Neetika Bansal, Dr. Vishal Goyal, Dr. Simpel Rani

Code mixing is prevalent when users use two or more languages while communicating. It becomes more complex when users prefer romanized text to Unicode typing. The automatic processing of social media data has become one of popular areas of interest. Especially since COVID period the involvement of youngsters has attained heights. Walking with the pace our intended software deals with Language Identification and Normalization of English and Punjabi code mixed text. The software designed follows a pipeline which includes data collection, pre-processing, language identification, handling Out of Vocabulary words, normalization and transliteration of English- Punjabi text. After applying five-fold cross validation on the corpus, the accuracy of 96.8% is achieved on a trained dataset of around 80025 tokens. After the prediction of the tags: the slangs, contractions in the user input are normalized to their standard form. In addition, the words with Punjabi as predicted tags are transliterated to Punjabi.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationTransliteration

Similar Papers 제목 키워드 기반

Normalization of Indonesian-English Code-Mixed Twitter Data

2019-11-01 · WS 2019 11 · Anab Maulana Barik, Rahmad Mahendra, Mirna Adriani

Twitter is an excellent source of data for NLP researches as it offers tremendous amount of textual data. However, processing tweet to extract meaningful information is very challenging, at least for two reasons: (i) usi…

Language IdentificationLexical NormalizationTranslation

Sentiment Analysis in Code-Mixed Telugu-English Text with Unsupervised Data Normalization

2021-09-01 · RANLP 2021 9 · Siva Subrahamanyam Varma Kusampudi, Preetham Sathineni, Radhika Mamidi

In a multilingual society, people communicate in more than one language, leading to Code-Mixed data. Sentimental analysis on Code-Mixed Telugu-English Text (CMTET) poses unique challenges. The unstructured nature of the …

Sentiment Analysis

Speech Synthesis of Code-Mixed Text

2016-05-01 · LREC 2016 5 · Sunayana Sitaram, Alan W. black

Most Text to Speech (TTS) systems today assume that the input text is in a single language and is written in the same language that the text needs to be synthesized in. However, in bilingual and multilingual communities,…

Language IdentificationSpeech Synthesistext-to-speechText to Speech

MHE: Code-Mixed Corpora for Similar Language Identification

2022-06-01 · LREC 2022 6 · Priya Rani, John P. McCrae, Theodorus Fransen

This paper introduces a new Magahi-Hindi-English (MHE) code-mixed data-set for similar language identification (SMLID), where Magahi is a less-resourced minority language. This corpus provides a language id at two levels…

Language IdentificationSentence

DOSA: Dravidian Code-Mixed Offensive Span Identification Dataset

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Manikandan Ravikiran, Subbiah Annamalai

This paper presents the Dravidian Offensive Span Identification Dataset (DOSA) for under-resourced Tamil-English and Kannada-English code-mixed text. The dataset addresses the lack of code-mixed datasets with annotated o…

Language Identification