paper-with-me

홈 › Papers

Single-sample writers -- "Document Filter" and their impacts on writer identification

2020-05-18 · Fabio Pinhelli, Alceu S. Britto Jr, Luiz S. Oliveira, Yandre M. G. Costa, Diego Bertolini

The writing can be used as an important biometric modality which allows to unequivocally identify an individual. It happens because the writing of two different persons present differences that can be explored both in terms of graphometric properties or even by addressing the manuscript as a digital image, taking into account the use of image processing techniques that can properly capture different visual attributes of the image (e.g. texture). In this work, perform a detailed study in which we dissect whether or not the use of a database with only a single sample taken from some writers may skew the results obtained in the experimental protocol. In this sense, we propose here what we call "document filter". The "document filter" protocol is supposed to be used as a preprocessing technique, such a way that all the data taken from fragments of the same document must be placed either into the training or into the test set. The rationale behind it, is that the classifier must capture the features from the writer itself, and not features regarding other particularities which could affect the writing in a specific document (i.e. emotional state of the writer, pen used, paper type, and etc.). By analyzing the literature, one can find several works dealing the writer identification problem. However, the performance of the writer identification systems must be evaluated also taking into account the occurrence of writer volunteers who contributed with a single sample during the creation of the manuscript databases. To address the open issue investigated here, a comprehensive set of experiments was performed on the IAM, BFL and CVL databases. They have shown that, in the most extreme case, the recognition rate obtained using the "document filter" protocol drops from 81.80% to 50.37%.

📄 PDF Abstract BibTeX arXiv:2005.08424

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ICDAR 2019 Competition on Image Retrieval for Historical Handwritten Documents

2019-12-08 · Vincent Christlein, Anguelos Nicolaou, Mathias Seuret, Dominique Stutzmann 외

This competition investigates the performance of large-scale retrieval of historical document images based on writing style. Based on large image data sets provided by cultural heritage institutions and digital libraries…

Image RetrievalRetrievalWriter Retrieval

New Mathematical and Algorithmic Schemes for Pattern Classification with Application to the Identification of Writers of Important Ancient Documents

2013-06-28 · Dimitris Arabadjis, Fotios Giannopoulos, Constantin Papaodysseus, Solomon Zannos 외

In this paper, a novel approach is introduced for classifying curves into proper families, according to their similarity. First, a mathematical quantity we call plane curvature is introduced and a number of propositions …

General Classification

Towards Generating Citation Sentences for Multiple References with Intent Control

2021-12-02 · Jia-Yan Wu, Alexander Te-Wei Shieh, Shih-Ju Hsu, Yun-Nung Chen

Machine-generated citation sentences can aid automated scientific literature review and assist article writing. Current methods in generating citation text were limited to single citation generation using the citing docu…

DecoderSentence

Read, Revise, Repeat: A System Demonstration for Human-in-the-loop Iterative Text Revision

2022-04-07 · In2Writing (ACL) 2022 5 · Wanyu Du, Zae Myung Kim, Vipul Raheja, Dhruv Kumar 외

Revision is an essential part of the human writing process. It tends to be strategic, adaptive, and, more importantly, iterative in nature. Despite the success of large language models on text revision tasks, they are li…

VML-MOC: Segmenting a multiply oriented and curved handwritten text lines dataset

2021-01-19 · Berat Kurar Barakat, Rafi Cohen, Irina Rabaev, Jihad El-Sana

This paper publishes a natural and very complicated dataset of handwritten documents with multiply oriented and curved text lines, namely VML-MOC dataset. These text lines were written as remarks on the page margins by d…