Estimating Text Similarity based on Semantic Concept Embeddings
Due to their ease of use and high accuracy, Word2Vec (W2V) word embeddings enjoy great success in the semantic representation of words, sentences, and whole documents as well as for semantic similarity estimation. However, they have the shortcoming that they are directly extracted from a surface representation, which does not adequately represent human thought processes and also performs poorly for highly ambiguous words. Therefore, we propose Semantic Concept Embeddings (CE) based on the MultiNet Semantic Network (SN) formalism, which addresses both shortcomings. The evaluation on a marketing target group distribution task showed that the accuracy of predicted target groups can be increased by combining traditional word embeddings with semantic CEs.
Code (0)
등록된 구현이 없습니다.
Tasks
MarketingSemantic SimilaritySemantic Textual Similaritytext similarityWord EmbeddingsSimilar Papers 제목 키워드 기반
Estimating Mutual Information Between Dense Word Embeddings
Word embedding-based similarity measures are currently among the top-performing methods on unsupervised semantic textual similarity (STS) tasks. Recent work has increasingly adopted a statistical view on these embeddings…
Semantic Textual SimilaritySTSWord EmbeddingsOntology-Aware Token Embeddings for Prepositional Phrase Attachment
Type-level word embeddings use the same set of parameters to represent all instances of a word regardless of its context, ignoring the inherent lexical ambiguity in language. Instead, we embed semantic concepts (or synse…
Prepositional Phrase AttachmentWord EmbeddingsMedical Concept Normalization in User-Generated Texts by Learning Target Concept Embeddings
Medical concept normalization helps in discovering standard concepts in free-form text i.e., maps health-related mentions to standard concepts in a clinical knowledge base. It is much beyond simple string matching and re…
Clinical KnowledgeMedical Concept Normalizationtext-classificationText Classification+1Multi-Ontology Refined Embeddings (MORE): A Hybrid Multi-Ontology and Corpus-based Semantic Representation for Biomedical Concepts
Objective: Currently, a major limitation for natural language processing (NLP) analyses in clinical applications is that a concept can be referenced in various forms across different texts. This paper introduces Multi-On…
Word EmbeddingsMedical Concept Normalization in User Generated Texts by Learning Target Concept Embeddings
Medical concept normalization helps in discovering standard concepts in free-form text i.e., maps health-related mentions to standard concepts in a vocabulary. It is much beyond simple string matching and requires a deep…
General ClassificationMedical Concept Normalizationtext-classificationText Classification+1