paper-with-me

Papers

Informational Space of Meaning for Scientific Texts

2020-04-28 · Neslihan Suzen, Evgeny M. Mirkes, Alexander N. Gorban

In Natural Language Processing, automatic extracting the meaning of texts constitutes an important problem. Our focus is the computational analysis of meaning of short scientific texts (abstracts or brief reports). In this paper, a vector space model is developed for quantifying the meaning of words and texts. We introduce the Meaning Space, in which the meaning of a word is represented by a vector of Relative Information Gain (RIG) about the subject categories that the text belongs to, which can be obtained from observing the word in the text. This new approach is applied to construct the Meaning Space based on Leicester Scientific Corpus (LSC) and Leicester Scientific Dictionary-Core (LScDC). The LSC is a scientific corpus of 1,673,350 abstracts and the LScDC is a scientific dictionary which words are extracted from the LSC. Each text in the LSC belongs to at least one of 252 subject categories of Web of Science (WoS). These categories are used in construction of vectors of information gains. The Meaning Space is described and statistically analysed for the LSC with the LScDC. The usefulness of the proposed representation model is evaluated through top-ranked words in each category. The most informative n words are ordered. We demonstrated that RIG-based word ranking is much more useful than ranking based on raw word frequency in determining the science-specific meaning and importance of a word. The proposed model based on RIG is shown to have ability to stand out topic-specific words in categories. The most informative words are presented for 252 categories. The new scientific dictionary and the 103,998 x 252 Word-Category RIG Matrix are available online. Analysis of the Meaning Space provides us with a tool to further explore quantifying the meaning of a text using more complex and context-dependent meaning models that use co-occurrence of words and their combinations.

📄 PDF Abstract BibTeX arXiv:2004.13717

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Informational Space Based Semantic Analysis for Scientific Texts

2022-05-31 · Neslihan Suzen, Alexander N. Gorban, Jeremy Levesley, Evgeny M. Mirkes

One major problem in Natural Language Processing is the automatic analysis and representation of human language. Human language is ambiguous and deeper understanding of semantics and creating human-to-machine interaction…

Common Sense Reasoning

Semantic Analysis for Automated Evaluation of the Potential Impact of Research Articles

2021-04-26 · Neslihan Suzen, Alexander Gorban, Jeremy Levesley, Evgeny Mirkes

Can the analysis of the semantics of words used in the text of a scientific paper predict its future impact measured by citations? This study details examples of automated text classification that achieved 80% success ra…

ArticlesCitation PredictionGeneral Classificationtext-classification+1

Principal Components of the Meaning

2020-09-18 · Neslihan Suzen, Alexander Gorban, Jeremy Levesley, Evgeny Mirkes

In this paper we argue that (lexical) meaning in science can be represented in a 13 dimension Meaning Space. This space is constructed using principal component analysis (singular decomposition) on the matrix of word cat…

The Underlying Dynamics of Life and Its Evolution: A Prigogine-Inspired Informational Dissipative System

2024-12-03 · Salvatore Chirumbolo, Antonio Vella

Life is fundamentally a scientific enigma. The interplay between chaos, entropy dynamics, and Prigogine's dissipative systems offers profound insights into the emergence, stabilization, and eventual collapse of far-from-…

Estimating Linguistic Complexity for Science Texts

2018-06-01 · WS 2018 6 · Farah Nadeem, Mari Ostendorf

Evaluation of text difficulty is important both for downstream tasks like text simplification, and for supporting educators in classrooms. Existing work on automated text complexity analysis uses linear models with engin…

Feature EngineeringReading ComprehensionText Simplification