paper-with-me

홈 › Papers

Beyond String Matching: Semantic Evaluation of PDF Table Extraction

2026-03-19 · Pius Horn, Janis Keuper arxiv

Reliably extracting tables from PDFs is essential for large-scale scientific data mining and knowledge base construction, yet existing evaluation approaches rely on rule-based metrics that fail to capture semantic equivalence of table content. We present a benchmarking framework based on synthetically generated PDFs with precise LaTeX ground truth, using tables sourced from arXiv to ensure realistic complexity and diversity. As our central methodological contribution, we apply LLM-as-a-judge for semantic table evaluation, integrated into a matching pipeline that accommodates inconsistencies in parser outputs. Through a human validation study comprising over 1,500 quality judgments on extracted table pairs, we show that LLM-based evaluation achieves substantially higher correlation with human judgment (Pearson r=0.93) compared to currently used Tree Edit Distance-based Similarity (TEDS, r=0.68) and Grid Table Similarity (GriTS, r=0.70). Evaluating 21 contemporary PDF parsers across 100 synthetic documents containing 451 tables reveals significant performance disparities. Our results offer practical guidance for selecting parsers for tabular data extraction and establish a reproducible, scalable evaluation methodology for this critical task. Code and data: https://github.com/phorn1/pdf-parse-bench Metric study and human evaluation: https://github.com/phorn1/table-metric-study

📄 PDF Abstract BibTeX arXiv:2603.18652

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bridging Classical and Quantum String Matching: A Computational Reformulation of Bit-Parallelism

2025-03-07 · Simone Faro, Arianna Pavone, Caterina Viola

String matching is a fundamental problem in computer science, with critical applications in text retrieval, bioinformatics, and data analysis. Among the numerous solutions that have emerged for this problem in recent dec…

Text Retrieval

Permutation Matching Under Parikh Budgets: Linear-Time Detection, Packing, and Disjoint Selection

2026-01-14 · MD Nazmul Alam Shanto, Md. Tanzeem Rahat, Md. Manzurul Hasan arxiv

We study permutation (jumbled/Abelian) pattern matching over a general alphabet $Σ$. Given a pattern P of length m and a text T of length n, the classical task is to decide whether T contains a length-m substring whose P…

SEMA: Text Simplification Evaluation through Semantic Alignment

2020-12-01 · AACL (NLP-TEA) 2020 12 · Xuan Zhang, Huizhou Zhao, Kexin Zhang, Yiyang Zhang

Text simplification is an important branch of natural language processing. At present, methods used to evaluate the semantic retention of text simplification are mostly based on string matching. We propose the SEMA (text…

Text Simplification

Scalable Approach for Normalizing E-commerce Text Attributes (SANTA)

2021-06-12 · ACL (ECNLP) 2021 8 · Ravi Shankar Mishra, Kartik Mehta, Nikhil Rasiwasia

In this paper, we present SANTA, a scalable framework to automatically normalize E-commerce attribute values (e.g. "Win 10 Pro") to a fixed set of pre-defined canonical values (e.g. "Windows 10"). Earlier works on attrib…

AttributeTripletWord Similarity

Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents

2026-02-11 · Yifei Li, Weidong Guo, Lingling Zhang, Rongman Xu 외 arxiv

Long-term conversational memory is a core capability for LLM-based dialogue systems, yet existing benchmarks and evaluation protocols primarily focus on surface-level factual recall. In realistic interactions, appropriat…