paper-with-me

Papers

Unsupervised document zone identification using probabilistic graphical models

2012-05-01 · LREC 2012 5 · Andrea Varga, Daniel Preo{\c{t}}iuc-Pietro, Fabio Ciravegna

Document zone identification aims to automatically classify sequences of text-spans (e.g. sentences) within a document into predefined zone categories. Current approaches to document zone identification mostly rely on supervised machine learning methods, which require a large amount of annotated data, which is often difficult and expensive to obtain. In order to overcome this bottleneck, we propose graphical models based on the popular Latent Dirichlet Allocation (LDA) model. The first model, which we call zoneLDA aims to cluster the sentences into zone classes using only unlabelled data. We also study an extension of zoneLDA called zoneLDAb, which makes distinction between common words and non-common words within the different zone types. We present results on two different domains: the scientific domain and the technical domain. For the latter one we propose a new document zone classification schema, which has been annotated over a collection of 689 documents, achieving a Kappa score of 85{\%}. Overall our experiments show promising results for both of the domains, outperforming the baseline model. Furthermore, on the technical domain the performance of the models are comparable to the supervised approach using the same feature sets. We thus believe that graphical models are a promising avenue of research for automatic document zoning.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bilbo-Val: Automatic Identification of Bibliographical Zone in Papers

2016-05-01 · LREC 2016 5 · Amal Htait, Sebastien Fournier, Patrice Bellot

In this paper, we present the automatic annotation of bibliographical references{'} zone in papers and articles of XML/TEI format. Our work is applied through two phases: first, we use machine learning technology to clas…

Articles

SGFusion: Stochastic Geographic Gradient Fusion in Federated Learning

2025-10-27 · Khoa Nguyen, Khang Tran, NhatHai Phan, Cristian Borcea 외 arxiv

This paper proposes Stochastic Geographic Gradient Fusion (SGFusion), a novel training algorithm to leverage the geographic information of mobile users in Federated Learning (FL). SGFusion maps the data collected by mobi…

Federated Learning

Automatic identification of document sections for designing a French clinical corpus (Identification automatique de zones dans des documents pour la constitution d'un corpus m\'edical en fran\ccais) [in French]

2014-07-01 · JEPTALNRECITAL 2014 7 · Louise Del{\'e}ger, Aur{\'e}lie N{\'e}v{\'e}ol

A Probabilistic Generative Model for Typographical Analysis of Early Modern Printing

2020-05-04 · ACL 2020 6 · Kartik Goyal, Chris Dyer, Christopher Warren, Max G'Sell 외

We propose a deep and interpretable probabilistic generative model to analyze glyph shapes in printed Early Modern documents. We focus on clustering extracted glyph images into underlying templates in the presence of mul…

Clustering

Unsupervised Text Deidentification

2022-10-20 · John X. Morris, Justin T. Chiu, Ramin Zabih, Alexander M. Rush

Deidentification seeks to anonymize textual data prior to distribution. Automatic deidentification primarily uses supervised named entity recognition from human-labeled data points. We propose an unsupervised deidentific…

Named Entity RecognitionNamed Entity Recognition (NER)