paper-with-me

Papers Table Extraction

“Table Extraction” 태그가 달린 논문 32편 · 필터 해제

Design and Implementation of an OCR-Powered Pipeline for Table Extraction from Invoices

2025-07-09 · Parshva Dhilankumar Patel

This paper presents the design and development of an OCR-powered pipeline for efficient table extraction from invoices. The system leverages Tesseract OCR for text recognition and custom post-processing logic to detect, …

Boundary DetectionOptical Character Recognition (OCR)Table Extraction

QUEST: Quality-aware Semi-supervised Table Extraction for Business Documents

2025-06-17 · Eliott Thomas, Mickael Coustaty, Aurelie Joseph, Gaspar Deloin 외

Automating table extraction (TE) from business documents is critical for industrial workflows but remains challenging due to sparse annotations and error-prone multi-stage pipelines. While semi-supervised learning (SSL) …

Pseudo LabelTable Extraction

RAPTOR: Refined Approach for Product Table Object Recognition

2025-02-19 · Eliott Thomas, Mickael Coustaty, Aurelie Joseph, Gaspar Deloin 외

Extracting tables from documents is a critical task across various industries, especially on business documents like invoices and reports. Existing systems based on DEtection TRansformer (DETR) such as TAble TRansformer …

ObjectObject RecognitionTable DetectionTable Extraction

ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution

2025-02-03 · Kanika Goswami, Puneet Mathur, Ryan Rossi, Franck Dernoncourt

Large Language Models (LLMs) can perform chart question-answering tasks but often generate unverified hallucinated responses. Existing answer attribution methods struggle to ground responses in source charts due to limit…

Chart Question AnsweringQuestion AnsweringRe-RankingRetrieval+1

CISOL: An Open and Extensible Dataset for Table Structure Recognition in the Construction Industry

2025-01-26 · David Tschirschwitz, Volker Rodehorst

Reproducibility and replicability are critical pillars of empirical research, particularly in machine learning, where they depend not only on the availability of models, but also on the datasets used to train and evaluat…

BenchmarkingObject DetectionTable DetectionTable Extraction

PlotEdit: Natural Language-Driven Accessible Chart Editing in PDFs via Multimodal LLM Agents

2025-01-20 · Kanika Goswami, Puneet Mathur, Ryan Rossi, Franck Dernoncourt

Chart visualizations, while essential for data interpretation and communication, are predominantly accessible only as images in PDFs, lacking source data tables and stylistic information. To enable effective editing of c…

AttributeTable Extraction

SynFinTabs: A Dataset of Synthetic Financial Tables for Information and Table Extraction

2024-12-05 · Ethan Bradley, Muhammad Roman, Karen Rafferty, Barry Devereux

Table extraction from document images is a challenging AI problem, and labelled data for many content domains is difficult to come by. Existing table extraction datasets often focus on scientific tables due to the vast a…

ArticlesDataset GenerationExtractive Question-AnsweringLarge Language Model+3

PdfTable: A Unified Toolkit for Deep Learning-Based Table Extraction

2024-09-08 · Lei Sheng, Shuai-Shuai Xu

Currently, a substantial volume of document data exists in an unstructured format, encompassing Portable Document Format (PDF) files and images. Extracting information from these documents presents formidable challenges …

Deep LearningDocument Layout AnalysisOptical Character RecognitionOptical Character Recognition (OCR)+2

tabulapdf: An R Package to Extract Tables from PDF Documents

2024-08-25 · Mauricio Vargas Sepúlveda, Thomas J. Leeper, Tom Paskhalis, Manuel Aristarán 외

tabulapdf is an R package that utilizes the Tabula Java library to import tables from PDF files directly into R. This tool can reduce time and effort in data extraction processes in fields like investigative journalism. …

RetrievalTable Extraction

H-STAR: LLM-driven Hybrid SQL-Text Adaptive Reasoning on Tables

2024-06-29 · Nikhil Abhyankar, Vivek Gupta, Dan Roth, Chandan K. Reddy

Tabular reasoning involves interpreting natural language queries about tabular data, which presents a unique challenge of combining language understanding with structured data analysis. Existing methods employ either tex…

Fact VerificationMathematical ReasoningNatural Language QueriesQuestion Answering+2

Financial Table Extraction in Image Documents

2024-03-18 · William Watson, Bo Liu

Table extraction has long been a pervasive problem in financial services. This is more challenging in the image domain, where content is locked behind cumbersome pixel format. Luckily, advances in deep learning for image…

Image SegmentationOptical Character Recognition (OCR)Semantic SegmentationTable Extraction

SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials

2024-02-22 · Wonjoong Kim, Sangwu Park, Yeonjun In, Seokwon Han 외

Recently, interpreting complex charts with logical reasoning has emerged as challenges due to the development of vision-language models. A prior state-of-the-art (SOTA) model has presented an end-to-end method that lever…

Chart Question AnsweringLanguage ModelingLanguage ModellingLarge Language Model+3

Schema-Driven Information Extraction from Heterogeneous Tables

2023-05-23 · Fan Bai, Junmo Kang, Gabriel Stanovsky, Dayne Freitag 외

In this paper, we explore the question of whether large language models can support cost-efficient information extraction from tables. We introduce schema-driven information extraction, a new task that transforms tabular…

Attribute ExtractionInstruction FollowingTable Extraction

A Benchmark of PDF Information Extraction Tools using a Multi-Task and Multi-Domain Evaluation Framework for Academic Documents

2023-03-17 · Norman Meuschke, Apurva Jagdale, Timo Spinde, Jelena Mitrović 외

Extracting information from academic PDF documents is crucial for numerous indexing, retrieval, and analysis use cases. Choosing the best tool to extract specific content elements is difficult because many, technically d…

RetrievalTable Extraction

CTE: A Dataset for Contextualized Table Extraction

2023-02-02 · Andrea Gemelli, Emanuele Vivoli, Simone Marinai

Relevant information in documents is often summarized in tables, helping the reader to identify useful facts. Most benchmark datasets support either document layout analysis or table understanding, but lack in providing …

Document Layout AnalysisTable DetectionTable Extraction

Deep learning for table detection and structure recognition: A survey

2022-11-15 · Mahmoud Kasem, Abdelrahman Abdallah, Alexander Berendeyev, Ebrahem Elkady 외

Tables are everywhere, from scientific journals, papers, websites, and newspapers all the way to items we buy at the supermarket. Detecting them is thus of utmost importance to automatically understanding the content of …

Deep Learningobject-detectionObject DetectionTable Detection+2

A two-stage approach for table extraction in invoices

2022-10-10 · Thomas Saout, Frédéric Lardeux, Frédéric Saubion

The automated analysis of administrative documents is an important field in document recognition that is studied for decades. Invoices are key documents among these huge amounts of documents available in companies and pu…

Table ExtractionVocal Bursts Valence Prediction

Graph Neural Networks and Representation Embedding for Table Extraction in PDF Documents

2022-08-23 · Andrea Gemelli, Emanuele Vivoli, Simone Marinai

Tables are widely used in several types of documents since they can bring important information in a structured way. In scientific papers, tables can sum up novel discoveries and summarize experimental results, making th…

Optical Character Recognition (OCR)Table Extraction

DiSCoMaT: Distantly Supervised Composition Extraction from Tables in Materials Science Articles

2022-07-03 · Tanishq Gupta, Mohd Zaki, Devanshi Khatsuriya, Kausik Hira 외

A crucial component in the curation of KB for a scientific domain (e.g., materials science, foods & nutrition, fuels) is information extraction from tables in the domain's published research articles. To facilitate resea…

ArticlesNutritionTable Extraction

Modelling the semantics of text in complex document layouts using graph transformer networks

2022-02-18 · Thomas Roland Barillot, Jacob Saks, Polena Lilyanova, Edward Torgas 외

Representing structured text from complex documents typically calls for different machine learning techniques, such as language models for paragraphs and convolutional neural networks (CNNs) for table extraction, which p…

Table Extraction
1–20 / 32 다음 →