paper-with-me

홈 › Papers

Text sampling strategies for predicting missing bibliographic links

2023-01-04 · F. V. Krasnova, I. S. Smaznevicha, E. N. Baskakova

The paper proposes various strategies for sampling text data when performing automatic sentence classification for the purpose of detecting missing bibliographic links. We construct samples based on sentences as semantic units of the text and add their immediate context which consists of several neighboring sentences. We examine a number of sampling strategies that differ in context size and position. The experiment is carried out on the collection of STEM scientific papers. Including the context of sentences into samples improves the result of their classification. We automatically determine the optimal sampling strategy for a given text collection by implementing an ensemble voting when classifying the same data sampled in different ways. Sampling strategy taking into account the sentence context with hard voting procedure leads to the classification accuracy of 98% (F1-score). This method of detecting missing bibliographic links can be used in recommendation engines of applied intelligent information systems.

📄 PDF Abstract BibTeX arXiv:2301.01673

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationPositionSentenceSentence Classification

Similar Papers 제목 키워드 기반

The Fellowship of the Authors: Disambiguating Names from Social Network Context

2022-08-31 · Ryan Muther, David Smith

Most NLP approaches to entity linking and coreference resolution focus on retrieving similar mentions using sparse or dense text representations. The common "Wikification" task, for instance, retrieves candidate Wikipedi…

Articlescoreference-resolutionCoreference ResolutionEntity Linking+1

A natural language interface to a graph-based bibliographic information retrieval system

2016-12-10 · Yongjun Zhu, Erjia Yan, Il-Yeol Song

With the ever-increasing scientific literature, there is a need on a natural language interface to bibliographic information retrieval systems to retrieve related information effectively. In this paper, we propose a natu…

Information Retrievalnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2

Predicting Performance of Software Configurations: There is no Silver Bullet

2019-11-28 · Alexander Grebhahn, Norbert Siegmund, Sven Apel

Many software systems offer configuration options to tailor their functionality and non-functional properties (e.g., performance). Often, users are interested in the (performance-)optimal configuration, but struggle to f…

BIG-bench Machine LearningPrediction

Flexible variable selection in the presence of missing data

2022-02-25 · B. D. Williamson, Y. Huang

In many applications, it is of interest to identify a parsimonious set of features, or panel, from multiple candidates that achieves a desired level of performance in predicting a response. This task is often complicated…

ImputationVariable Selection

Bilbo-Val: Automatic Identification of Bibliographical Zone in Papers

2016-05-01 · LREC 2016 5 · Amal Htait, Sebastien Fournier, Patrice Bellot

In this paper, we present the automatic annotation of bibliographical references{'} zone in papers and articles of XML/TEI format. Our work is applied through two phases: first, we use machine learning technology to clas…

Articles