paper-with-me

홈 › Papers

A document processing pipeline for the construction of a dataset for topic modeling based on the judgments of the Italian Supreme Court

2025-05-13 · Matteo Marulli, Glauco Panattoni, Marco Bertini

Topic modeling in Italian legal research is hindered by the lack of public datasets, limiting the analysis of legal themes in Supreme Court judgments. To address this, we developed a document processing pipeline that produces an anonymized dataset optimized for topic modeling. The pipeline integrates document layout analysis (YOLOv8x), optical character recognition, and text anonymization. The DLA module achieved a mAP@50 of 0.964 and a mAP@50-95 of 0.800. The OCR detector reached a mAP@50-95 of 0.9022, and the text recognizer (TrOCR) obtained a character error rate of 0.0047 and a word error rate of 0.0248. Compared to OCR-only methods, our dataset improved topic modeling with a diversity score of 0.6198 and a coherence score of 0.6638. We applied BERTopic to extract topics and used large language models to generate labels and summaries. Outputs were evaluated against domain expert interpretations. Claude Sonnet 3.7 achieved a BERTScore F1 of 0.8119 for labeling and 0.9130 for summarization.

📄 PDF Abstract BibTeX arXiv:2505.08439

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityDocument Layout AnalysisOptical Character RecognitionOptical Character Recognition (OCR)Text Anonymization

Methods 이 논문이 사용한 방법론

DLA 설명 없음

Similar Papers 제목 키워드 기반

RaV-IDP: A Reconstruction-as-Validation Framework for Faithful Intelligent Document Processing

2026-04-26 · Pritesh Jha arxiv

Intelligent document processing pipelines extract structured entities (tables, images, and text) from documents for use in downstream systems such as knowledge bases, retrieval-augmented generation, and analytics. A pers…

Capturing research literature attitude towards Sustainable Development Goals: an LLM-based topic modeling approach

2024-11-05 · Francesco Invernici, Francesca Curati, Jelena Jakimov, Amirhossein Samavi 외

The world is facing a multitude of challenges that hinder the development of human civilization and the well-being of humanity on the planet. The Sustainable Development Goals (SDGs) were formulated by the United Nations…

HADES: Homologous Automated Document Exploration and Summarization

2023-02-25 · Piotr Wilczyński, Artur Żółkowski, Mateusz Krzyziński, Emilia Wiśnios 외

This paper introduces HADES, a novel tool for automatic comparative documents with similar structures. HADES is designed to streamline the work of professionals dealing with large volumes of documents, such as policy doc…

Neural Topic Model with Reinforcement Learning

2019-11-01 · IJCNLP 2019 11 · Lin Gui, Jia Leng, Gabriele Pergola, Yu Zhou 외

In recent years, advances in neural variational inference have achieved many successes in text processing. Examples include neural topic models which are typically built upon variational autoencoder (VAE) with an objecti…

modelreinforcement-learningReinforcement LearningReinforcement Learning (RL)+2

MAPEX: A Multi-Agent Pipeline for Keyphrase Extraction

2025-09-23 · Liting Zhang, Shiwan Zhao, Aobo Kong, Qicheng Li arxiv

Keyphrase extraction is a fundamental task in natural language processing. However, existing unsupervised prompt-based methods for Large Language Models (LLMs) often rely on single-stage inference pipelines with uniform …

Keyphrase Extraction