paper-with-me

홈 › Papers

A Simple and Efficient Probabilistic Language model for Code-Mixed Text

2021-06-29 · M Zeeshan Ansari, Tanvir Ahmad, M M Sufyan Beg, Asma Ikram

The conventional natural language processing approaches are not accustomed to the social media text due to colloquial discourse and non-homogeneous characteristics. Significantly, the language identification in a multilingual document is ascertained to be a preceding subtask in several information extraction applications such as information retrieval, named entity recognition, relation extraction, etc. The problem is often more challenging in code-mixed documents wherein foreign languages words are drawn into base language while framing the text. The word embeddings are powerful language modeling tools for representation of text documents useful in obtaining similarity between words or documents. We present a simple probabilistic approach for building efficient word embedding for code-mixed text and exemplifying it over language identification of Hindi-English short test messages scrapped from Twitter. We examine its efficacy for the classification task using bidirectional LSTMs and SVMs and observe its improved scores over various existing code-mixed embeddings

📄 PDF Abstract BibTeX arXiv:2106.15102

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalLanguage IdentificationLanguage ModelingLanguage Modellingnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Relation ExtractionRetrievalWord Embeddings

Similar Papers 제목 키워드 기반

A Fast, Compact, Accurate Model for Language Identification of Codemixed Text

2018-10-09 · EMNLP 2018 10 · Yuan Zhang, Jason Riesa, Daniel Gillick, Anton Bakalov 외

We address fine-grained multilingual language identification: providing a language code for every token in a sentence, including codemixed text containing multiple languages. Such text is prevalent online, in documents, …

DecoderLanguage IdentificationSentence

HCMS at SemEval-2020 Task 9: A Neural Approach to Sentiment Analysis for Code-Mixed Texts

2020-07-23 · SEMEVAL 2020 · Aditya Srivastava, V. Harsha Vardhan

Problems involving code-mixed language are often plagued by a lack of resources and an absence of materials to perform sophisticated transfer learning with. In this paper we describe our submission to the Sentimix Hindi-…

Sentiment AnalysisSentiment ClassificationText Classification

CodemixedNLP: An Extensible and Open NLP Toolkit for Code-Mixing

2021-06-10 · NAACL (CALCS) 2021 6 · Sai Muralidhar Jayanthi, Kavya Nerella, Khyathi Raghavi Chandu, Alan W Black

The NLP community has witnessed steep progress in a variety of tasks across the realms of monolingual and multilingual language processing recently. These successes, in conjunction with the proliferating mixed language i…

An Ensemble Model for Sentiment Analysis of Hindi-English Code-Mixed Data

2018-06-12 · Madan Gopal Jhanwar, Arpita Das

In multilingual societies like India, code-mixed social media texts comprise the majority of the Internet. Detecting the sentiment of the code-mixed user opinions plays a crucial role in understanding social, economic an…

Sentiment Analysis

Exploring Text-to-Text Transformers for English to Hinglish Machine Translation with Synthetic Code-Mixing

2021-05-18 · NAACL (CALCS) 2021 6 · Ganesh Jawahar, El Moatez Billah Nagoudi, Muhammad Abdul-Mageed, Laks V. S. Lakshmanan

We describe models focused at the understudied problem of translating between monolingual and code-mixed language pairs. More specifically, we offer a wide range of models that convert monolingual English text into Hingl…

DecoderLanguage ModellingMachine TranslationTranslation