paper-with-me

Papers

Reference String Extraction Using Line-Based Conditional Random Fields

2017-05-23 · Körner Martin

The extraction of individual reference strings from the reference section of scientific publications is an important step in the citation extraction pipeline. Current approaches divide this task into two steps by first detecting the reference section areas and then grouping the text lines in such areas into reference strings. We propose a classification model that considers every line in a publication as a potential part of a reference string. By applying line-based conditional random fields rather than constructing the graphical model based on the individual words, dependencies and patterns that are typical in reference sections provide strong features while the overall complexity of the model is reduced.

📄 PDF Abstract BibTeX arXiv:1705.08154

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Synthetic vs. Real Reference Strings for Citation Parsing, and the Importance of Re-training and Out-Of-Sample Data for Meaningful Evaluations: Experiments with GROBID, GIANT and Cora

2020-04-22 · WOSP 2020 8 · Mark Grennan, Joeran Beel

Citation parsing, particularly with deep neural networks, suffers from a lack of training data as available datasets typically contain only a few thousand training instances. Manually labelling citation strings is very t…

A Linguistic Model for Terminology Extraction based Conditional Random Fields

2012-09-30 · Fethi Fkih, Mohamed Nazih Omri, Imen Toumia

In this paper, we show the possibility of using a linear Conditional Random Fields (CRF) for terminology extraction from a specialized text corpus.

From "Strings" to "Things" for Personal Knowledge Graphs: Evaluating LLM Triple Extraction for Recommendation Systems

2026-04-18 · Abhirup Dasgupta, Fernando Spadea, Oshani Seneviratne arxiv

Personal Knowledge Graphs (PKGs) offer a privacy-preserving framework for modeling user preferences, yet constructing them from unstructured, decentralized conversational data remains a challenge. This paper bridges the …

Recommendation SystemsKnowledge Graphs

Using BibTeX to Automatically Generate Labeled Data for Citation Field Extraction

2020-06-09 · AKBC 2020 6 · Dung Thai, Zhiyang Xu, Nicholas Monath, Boris Veytsman 외

Accurate parsing of citation reference strings is crucial to automatically construct scholarly databases such as Google Scholar or Semantic Scholar. Citation field extraction (CFE) is precisely this task---given a refere…

Management

Bundesrecht: An Open Library and Corpus for German Statutory Reference Processing

2026-05-29 · Harshil Darji, Martin Heckelmann, Christina Kratsch, Gerard de Melo arxiv

Statutory references are central to legal language understanding, but are difficult to process automatically, as they appear in compact and variable surface forms, may combine multiple targets, use special abbreviations,…

Information Extraction