paper-with-me

홈 › Papers

Analysis Of Contextual and Non-Contextual Word Embedding Models For Hindi NER With Web Application For Data Collection

2021-02-18 · 10th International Advanced Computing Conference 2021 2 · Aindriya Barua, Thara.S, Premjith B, Soman KP‡

Named Entity Recognition (NER) is the process of taking a string and identifying relevant proper nouns in it. In this paper ‡ we report the development of the Hindi NER system, in Devanagari script, using various embedding models. We categorize embeddings as Contextual and Non-contextual, and further compare them inter and intra-category. Under non-contextual type embeddings, we experiment with Word2Vec and FastText, and under the contextual embedding category, we experiment with BERT and its variants, viz. RoBERTa, ELECTRA, CamemBERT, Distil-BERT, XLM-RoBERTa. For non-contextual embeddings, we use five machine learning algorithms namely Gaussian NB, Adaboost Classifier, Multi-layer Perceptron classifier, Random Forest Classifier, and Decision Tree Classifier for developing ten Hindi NER systems, each, once with Fast Text and once with Gensim Word2Vec word embedding models. These models are then compared with Transformers based contextual NER models, using BERT and its variants. A comparative study among all these NER models is made. Finally, the best of all these models is used and a web app is built, that takes a Hindi text of any length and returns NER tags for each word and takes feedback from the user about the correctness of tags. These feed-backs aid our further data collection.

📄 PDF Abstract BibTeX

Code (1)

AindriyaBarua/Contextual-vs-Non-Contextual-Word-Embeddings-For-Hindi-NER-With-WebApp

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
WordPiece 설명 없음
Adam 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
BERT BERT, or Bidirectional Encoder Representations from Transformers, improves upon standard Transformers by removing the…

Similar Papers 제목 키워드 기반

HinFlair: pre-trained contextual string embeddings for pos tagging and text classification in the Hindi language

2021-01-18 · Harsh Patel

Recent advancements in language models based on recurrent neural networks and transformers architecture have achieved state-of-the-art results on a wide range of natural language processing tasks such as pos tagging, nam…

ClassificationGeneral Classificationnamed-entity-recognitionNamed Entity Recognition+5

Contextual Mood Analysis with Knowledge Graph Representation for Hindi Song Lyrics in Devanagari Script

2021-08-16 · Makarand Velankar, Rachita Kotian, Parag Kulkarni

Lyrics play a significant role in conveying the song's mood and are information to understand and interpret music communication. Conventional natural language processing approaches use translation of the Hindi text into …

Incremental LearningRetrievalTranslation

Multilingual Offensive Language Identification with Cross-lingual Embeddings

2020-10-11 · EMNLP 2020 11 · Tharindu Ranasinghe, Marcos Zampieri

Offensive content is pervasive in social media and a reason for concern to companies and government organizations. Several studies have been recently published investigating methods to detect the various forms of such co…

Language IdentificationTransfer LearningWord Embeddings

Context based Analysis of Lexical Semantics for Hindi Language

2019-01-23 · Mohd Zeeshan Ansari, Lubna Khan

A word having multiple senses in a text introduces the lexical semantic task to find out which particular sense is appropriate for the given context. One such task is Word sense disambiguation which refers to the identif…

Word Sense Disambiguation

"A Passage to India": Pre-trained Word Embeddings for Indian Languages

2021-12-27 · Kumar Saurav, Kumar Saunack, Diptesh Kanojia, Pushpak Bhattacharyya

Dense word vectors or 'word embeddings' which encode semantic properties of words, have now become integral to NLP tasks like Machine Translation (MT), Question Answering (QA), Word Sense Disambiguation (WSD), and Inform…

Information RetrievalMachine TranslationNERQuestion Answering+3