Varying Vector Representations and Integrating Meaning Shifts into a PageRank Model for Automatic Term Extraction
We perform a comparative study for automatic term extraction from domain-specific language using a PageRank model with different edge-weighting methods. We vary vector space representations within the PageRank graph algorithm, and we go beyond standard co-occurrence and investigate the influence of measures of association strength and first- vs. second-order co-occurrence. In addition, we incorporate meaning shifts from general to domain-specific language as personalized vectors, in order to distinguish between termhood strengths of ambiguous words across word senses. Our study is performed for two domain-specific English corpora: ACL and do-it-yourself (DIY); and a domain-specific German corpus: cooking. The models are assessed by applying average precision and the roc score as evaluation metrices.
Code (0)
등록된 구현이 없습니다.
Tasks
Term ExtractionSimilar Papers 제목 키워드 기반
Controlling the Imprint of Passivization and Negation in Contextualized Representations
Contextualized word representations encode rich information about syntax and semantics, alongside specificities of each context of use. While contextual variation does not always reflect actual meaning shifts, it can sti…
Language ModelingLanguage ModellingMasked Language ModelingNegation+1Statistically Significant Detection of Linguistic Change
We propose a new computational approach for tracking and detecting statistically significant linguistic shifts in the meaning and usage of words. Such linguistic shifts are especially prevalent on the Internet, where the…
Change Point DetectionTime SeriesTime Series AnalysisSemantic Representations of Mathematical Expressions in a Continuous Vector Space
Mathematical notation makes up a large portion of STEM literature, yet finding semantic representations for formulae remains a challenging problem. Because mathematical notation is precise, and its meaning changes signif…
Training Temporal Word Embeddings with a Compass
Temporal word embeddings have been proposed to support the analysis of word meaning shifts during time and to study the evolution of languages. Different approaches have been proposed to generate vector representations o…
Diachronic Word EmbeddingsWord EmbeddingsUsing neural topic models to track context shifts of words: a case study of COVID-related terms before and after the lockdown in April 2020
This paper explores lexical meaning changes in a new dataset, which includes tweets from before and after the COVID-related lockdown in April 2020. We use this dataset to evaluate traditional and more recent unsupervised…
Language ModelingLanguage ModellingTopic Models