Papers Table Extraction
“Table Extraction” 태그가 달린 논문 32편 · 필터 해제
Design and Implementation of an OCR-Powered Pipeline for Table Extraction from Invoices
This paper presents the design and development of an OCR-powered pipeline for efficient table extraction from invoices. The system leverages Tesseract OCR for text recognition and custom post-processing logic to detect, …
Boundary DetectionOptical Character Recognition (OCR)Table ExtractionQUEST: Quality-aware Semi-supervised Table Extraction for Business Documents
Automating table extraction (TE) from business documents is critical for industrial workflows but remains challenging due to sparse annotations and error-prone multi-stage pipelines. While semi-supervised learning (SSL) …
Pseudo LabelTable ExtractionRAPTOR: Refined Approach for Product Table Object Recognition
Extracting tables from documents is a critical task across various industries, especially on business documents like invoices and reports. Existing systems based on DEtection TRansformer (DETR) such as TAble TRansformer …
ObjectObject RecognitionTable DetectionTable ExtractionChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution
Large Language Models (LLMs) can perform chart question-answering tasks but often generate unverified hallucinated responses. Existing answer attribution methods struggle to ground responses in source charts due to limit…
Chart Question AnsweringQuestion AnsweringRe-RankingRetrieval+1CISOL: An Open and Extensible Dataset for Table Structure Recognition in the Construction Industry
Reproducibility and replicability are critical pillars of empirical research, particularly in machine learning, where they depend not only on the availability of models, but also on the datasets used to train and evaluat…
BenchmarkingObject DetectionTable DetectionTable ExtractionPlotEdit: Natural Language-Driven Accessible Chart Editing in PDFs via Multimodal LLM Agents
Chart visualizations, while essential for data interpretation and communication, are predominantly accessible only as images in PDFs, lacking source data tables and stylistic information. To enable effective editing of c…
AttributeTable ExtractionSynFinTabs: A Dataset of Synthetic Financial Tables for Information and Table Extraction
Table extraction from document images is a challenging AI problem, and labelled data for many content domains is difficult to come by. Existing table extraction datasets often focus on scientific tables due to the vast a…
ArticlesDataset GenerationExtractive Question-AnsweringLarge Language Model+3PdfTable: A Unified Toolkit for Deep Learning-Based Table Extraction
Currently, a substantial volume of document data exists in an unstructured format, encompassing Portable Document Format (PDF) files and images. Extracting information from these documents presents formidable challenges …
Deep LearningDocument Layout AnalysisOptical Character RecognitionOptical Character Recognition (OCR)+2tabulapdf: An R Package to Extract Tables from PDF Documents
tabulapdf is an R package that utilizes the Tabula Java library to import tables from PDF files directly into R. This tool can reduce time and effort in data extraction processes in fields like investigative journalism. …
RetrievalTable ExtractionH-STAR: LLM-driven Hybrid SQL-Text Adaptive Reasoning on Tables
Tabular reasoning involves interpreting natural language queries about tabular data, which presents a unique challenge of combining language understanding with structured data analysis. Existing methods employ either tex…
Fact VerificationMathematical ReasoningNatural Language QueriesQuestion Answering+2Financial Table Extraction in Image Documents
Table extraction has long been a pervasive problem in financial services. This is more challenging in the image domain, where content is locked behind cumbersome pixel format. Luckily, advances in deep learning for image…
Image SegmentationOptical Character Recognition (OCR)Semantic SegmentationTable ExtractionSIMPLOT: Enhancing Chart Question Answering by Distilling Essentials
Recently, interpreting complex charts with logical reasoning has emerged as challenges due to the development of vision-language models. A prior state-of-the-art (SOTA) model has presented an end-to-end method that lever…
Chart Question AnsweringLanguage ModelingLanguage ModellingLarge Language Model+3Schema-Driven Information Extraction from Heterogeneous Tables
In this paper, we explore the question of whether large language models can support cost-efficient information extraction from tables. We introduce schema-driven information extraction, a new task that transforms tabular…
Attribute ExtractionInstruction FollowingTable ExtractionA Benchmark of PDF Information Extraction Tools using a Multi-Task and Multi-Domain Evaluation Framework for Academic Documents
Extracting information from academic PDF documents is crucial for numerous indexing, retrieval, and analysis use cases. Choosing the best tool to extract specific content elements is difficult because many, technically d…
RetrievalTable ExtractionCTE: A Dataset for Contextualized Table Extraction
Relevant information in documents is often summarized in tables, helping the reader to identify useful facts. Most benchmark datasets support either document layout analysis or table understanding, but lack in providing …
Document Layout AnalysisTable DetectionTable ExtractionDeep learning for table detection and structure recognition: A survey
Tables are everywhere, from scientific journals, papers, websites, and newspapers all the way to items we buy at the supermarket. Detecting them is thus of utmost importance to automatically understanding the content of …
Deep Learningobject-detectionObject DetectionTable Detection+2A two-stage approach for table extraction in invoices
The automated analysis of administrative documents is an important field in document recognition that is studied for decades. Invoices are key documents among these huge amounts of documents available in companies and pu…
Table ExtractionVocal Bursts Valence PredictionGraph Neural Networks and Representation Embedding for Table Extraction in PDF Documents
Tables are widely used in several types of documents since they can bring important information in a structured way. In scientific papers, tables can sum up novel discoveries and summarize experimental results, making th…
Optical Character Recognition (OCR)Table ExtractionDiSCoMaT: Distantly Supervised Composition Extraction from Tables in Materials Science Articles
A crucial component in the curation of KB for a scientific domain (e.g., materials science, foods & nutrition, fuels) is information extraction from tables in the domain's published research articles. To facilitate resea…
ArticlesNutritionTable ExtractionModelling the semantics of text in complex document layouts using graph transformer networks
Representing structured text from complex documents typically calls for different machine learning techniques, such as language models for paragraphs and convolutional neural networks (CNNs) for table extraction, which p…
Table Extraction