paper-with-me

Papers

EXmatcher: Combining Features Based on Reference Strings and Segments to Enhance Citation Matching

2019-06-11 · Behnam Ghavimi, Wolfgang Otto, Philipp Mayr

Citation matching is a challenging task due to different problems such as the variety of citation styles, mistakes in reference strings and the quality of identified reference segments. The classic citation matching configuration used in this paper is the combination of blocking technique and a binary classifier. Three different possible inputs (reference strings, reference segments and a combination of reference strings and segments) were tested to find the most efficient strategy for citation matching. In the classification step, we describe the effect which the probabilities of reference segments can have in citation matching. Our evaluation on a manually curated gold standard showed that the input data consisting of the combination of reference segments and reference strings lead to the best result. In addition, the usage of the probabilities of the segmentation slightly improves the result.

📄 PDF Abstract BibTeX arXiv:1906.04484

Code (0)

등록된 구현이 없습니다.

Tasks

BlockingGeneral Classification

Similar Papers 제목 키워드 기반

LexMatcher: Dictionary-centric Data Collection for LLM-based Machine Translation

2024-06-03 · Yongjing Yin, Jiali Zeng, Yafu Li, Fandong Meng 외

The fine-tuning of open-source large language models (LLMs) for machine translation has recently received considerable attention, marking a shift towards data-centric research from traditional neural machine translation.…

Data AugmentationMachine TranslationTranslationWord Sense Disambiguation

Reference String Extraction Using Line-Based Conditional Random Fields

2017-05-23 · Körner Martin

The extraction of individual reference strings from the reference section of scientific publications is an important step in the citation extraction pipeline. Current approaches divide this task into two steps by first d…

Synthetic vs. Real Reference Strings for Citation Parsing, and the Importance of Re-training and Out-Of-Sample Data for Meaningful Evaluations: Experiments with GROBID, GIANT and Cora

2020-04-22 · WOSP 2020 8 · Mark Grennan, Joeran Beel

Citation parsing, particularly with deep neural networks, suffers from a lack of training data as available datasets typically contain only a few thousand training instances. Manually labelling citation strings is very t…

Dependency Graph-to-String Statistical Machine Translation

2021-03-20 · Liangyou Li, Andy Way, Qun Liu

We present graph-based translation models which translate source graphs into target strings. Source graphs are constructed from dependency trees with extra links so that non-syntactic phrases are connected. Inspired by p…

Machine TranslationTranslation

Brightearth roads: Towards fully automatic road network extraction from satellite imagery

2024-06-21 · Liuyun Duan, Willard Mapurisa, Maxime Leras, Leigh Lotter 외

The modern road network topology comprises intricately designed structures that introduce complexity when automatically reconstructing road networks. While open resources like OpenStreetMap (OSM) offer road networks with…

Road Segmentation