Language Identification and Normalization of Code Mixed English and Punjabi Text
Code mixing is prevalent when users use two or more languages while communicating. It becomes more complex when users prefer romanized text to Unicode typing. The automatic processing of social media data has become one of popular areas of interest. Especially since COVID period the involvement of youngsters has attained heights. Walking with the pace our intended software deals with Language Identification and Normalization of English and Punjabi code mixed text. The software designed follows a pipeline which includes data collection, pre-processing, language identification, handling Out of Vocabulary words, normalization and transliteration of English- Punjabi text. After applying five-fold cross validation on the corpus, the accuracy of 96.8% is achieved on a trained dataset of around 80025 tokens. After the prediction of the tags: the slangs, contractions in the user input are normalized to their standard form. In addition, the words with Punjabi as predicted tags are transliterated to Punjabi.
Code (0)
등록된 구현이 없습니다.
Tasks
Language IdentificationTransliterationSimilar Papers 제목 키워드 기반
Normalization of Indonesian-English Code-Mixed Twitter Data
Twitter is an excellent source of data for NLP researches as it offers tremendous amount of textual data. However, processing tweet to extract meaningful information is very challenging, at least for two reasons: (i) usi…
Language IdentificationLexical NormalizationTranslationSentiment Analysis in Code-Mixed Telugu-English Text with Unsupervised Data Normalization
In a multilingual society, people communicate in more than one language, leading to Code-Mixed data. Sentimental analysis on Code-Mixed Telugu-English Text (CMTET) poses unique challenges. The unstructured nature of the …
Sentiment AnalysisSpeech Synthesis of Code-Mixed Text
Most Text to Speech (TTS) systems today assume that the input text is in a single language and is written in the same language that the text needs to be synthesized in. However, in bilingual and multilingual communities,…
Language IdentificationSpeech Synthesistext-to-speechText to SpeechMHE: Code-Mixed Corpora for Similar Language Identification
This paper introduces a new Magahi-Hindi-English (MHE) code-mixed data-set for similar language identification (SMLID), where Magahi is a less-resourced minority language. This corpus provides a language id at two levels…
Language IdentificationSentenceDOSA: Dravidian Code-Mixed Offensive Span Identification Dataset
This paper presents the Dravidian Offensive Span Identification Dataset (DOSA) for under-resourced Tamil-English and Kannada-English code-mixed text. The dataset addresses the lack of code-mixed datasets with annotated o…
Language Identification