Learning Contextual Embeddings for Structural Semantic Similarity using Categorical Information
Tree kernels (TKs) and neural networks are two effective approaches for automatic feature engineering. In this paper, we combine them by modeling context word similarity in semantic TKs. This way, the latter can operate subtree matching by applying neural-based similarity on tree lexical nodes. We study how to learn representations for the words in context such that TKs can exploit more focused information. We found that neural embeddings produced by current methods do not provide a suitable contextual similarity. Thus, we define a new approach based on a Siamese Network, which produces word representations while learning a binary text similarity. We set the latter considering examples in the same category as similar. The experiments on question and sentiment classification show that our semantic TK highly improves previous results.
Code (0)
등록된 구현이 없습니다.
Tasks
Feature EngineeringQuestion AnsweringRelation ExtractionSemantic SimilaritySemantic Textual SimilaritySentiment AnalysisSentiment Classificationtext similarityWord SimilarityMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Latte-Mix: Measuring Sentence Semantic Similarity with Latent Categorical Mixtures
Measuring sentence semantic similarity using pre-trained language models such as BERT generally yields unsatisfactory zero-shot performance, and one main reason is ineffective token aggregation methods such as mean pooli…
Semantic SimilaritySemantic Textual SimilaritySentenceSTS+1Contrastive Visual Semantic Pretraining Magnifies the Semantics of Natural Language Representations
We examine the effects of contrastive visual semantic pretraining by comparing the geometry and semantic properties of contextualized English language representations formed by GPT-2 and CLIP, a zero-shot multimodal imag…
Image CaptioningSemantic Textual SimilaritySentenceSentence Embeddings+1CitRet: A Hybrid Model for Cited Text Span Retrieval
The paper aims to identify cited text spans in the reference paper related to the given citance in the citing paper. We refer to it as cited text span retrieval (CTSR). Most current methods attempt this task by relying o…
RetrievalSemantic Textual SimilarityStructure Before Collapse: Transient semantic geometry in next-token prediction
Neural Collapse predicts that balanced one-hot classification pushes model representations to be equally far from each other; a symmetric configuration that depends only on the output label and ignores any semantic simil…
Semantic SimilarityEvaluating Word Embeddings with Categorical Modularity
We introduce categorical modularity, a novel low-resource intrinsic metric to evaluate word embedding quality. Categorical modularity is a graph modularity metric based on the $k$-nearest neighbor graph constructed with …
Bilingual Lexicon InductionSentiment AnalysisWord EmbeddingsWord Similarity