Papers Key Information Extraction
“Key Information Extraction” 태그가 달린 논문 92편 · 필터 해제
DocClaw: A Unified Agentic System for Intelligent Document Processing
Intelligent document processing (IDP) encompasses a broad range of tasks, including optical character recognition (OCR), document question answering (DocQA), and key information extraction (KIE). Despite their distinct o…
Key Information ExtractionQuestion AnsweringBeyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs
While the privacy risks of multimodal large language models (MLLMs) have drawn significant attention, the unique vulnerabilities of domain-specific MLLMs remain largely underexplored. Focusing on document understanding M…
Key Information ExtractionEnhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis
Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-server Large Multimodal Models (LMMs), wh…
Key Information ExtractionBaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension
Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing adoption in real-world, human-centric applications. However, this adoption is limit…
Key Information ExtractionDocument Layout AnalysisCC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing
Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capability in document literacy. However, their effectiveness in real-worl…
Key Information ExtractionQuestion AnsweringJaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
Japanese scene text poses challenges that multilingual benchmarks often fail to capture, including mixed scripts, frequent vertical writing, and a character inventory far larger than the Latin alphabet. Although Japanese…
Key Information ExtractionVisual Question AnsweringGLM-OCR Technical Report
GLM-OCR is an efficient 0.9B-parameter compact multimodal model designed for real-world document understanding. It combines a 0.4B-parameter CogViT visual encoder with a 0.5B-parameter GLM language decoder, achieving a s…
Key Information ExtractionComputational EfficiencyQianfan-OCR: A Unified End-to-End Model for Document Intelligence
We present Qianfan-OCR, a 4B-parameter end-to-end vision-language model that unifies document parsing, layout analysis, and document understanding within a single architecture. It performs direct image-to-Markdown conver…
Key Information ExtractionDocDjinn: Controllable Synthetic Document Generation with VLMs and Handwriting Diffusion
Effective document intelligence models rely on large amounts of annotated training data. However, procuring sufficient and high-quality data poses significant challenges due to the labor-intensive and costly nature of da…
Key Information ExtractionDocument Layout AnalysisDocument ClassificationQuestion AnsweringUNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
Key Information Extraction (KIE) from real-world documents remains challenging due to substantial variations in layout structures, visual quality, and task-specific information requirements. Recent Large Multimodal Model…
Key Information ExtractionUp to 36x Speedup: Mask-based Parallel Inference Paradigm for Key Information Extraction in MLLMs
Key Information Extraction (KIE) from visually-rich documents (VrDs) is a critical task, for which recent Large Language Models (LLMs) and Multi-Modal Large Language Models (MLLMs) have demonstrated strong potential. How…
Key Information ExtractionROAP: A Reading-Order and Attention-Prior Pipeline for Optimizing Layout Transformers in Key Information Extraction
The efficacy of Multimodal Transformers in visually-rich document understanding (VrDU) is critically constrained by two inherent limitations: the lack of explicit modeling for logical reading order and the interference o…
Key Information ExtractionMMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
Vision-language models enable the understanding and reasoning of complex traffic scenarios through multi-source information fusion, establishing it as a core technology for autonomous driving. However, existing vision-la…
Key Information ExtractionMultimodal ReasoningScene UnderstandingAutonomous DrivingFlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
Developing document understanding models at enterprise scale requires large, diverse, and well-annotated datasets spanning a wide range of document types. However, collecting such data is prohibitively expensive due to p…
Key Information ExtractionSynthetic Data GenerationSynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
Domain-specific Visually Rich Document Understanding (VRDU) presents significant challenges due to the complexity and sensitivity of documents in fields such as medicine, finance, and material science. Existing Large (Mu…
Key Information ExtractionSynthetic Data GenerationDomain AdaptationAgenticIE: An Adaptive Agent for Information Extraction from Complex Regulatory Documents
Declaration of Performance (DoP) documents, mandated by EU regulation, specify characteristics of construction products, such as fire resistance and insulation. While this information is essential for quality control and…
Key Information ExtractionQuestion AnsweringContext-Adaptive Synthesis and Compression for Enhanced Retrieval-Augmented Generation in Complex Domains
Large Language Models (LLMs) excel in language tasks but are prone to hallucinations and outdated knowledge. Retrieval-Augmented Generation (RAG) mitigates these by grounding LLMs in external knowledge. However, in compl…
Key Information ExtractionQuestion AnsweringVDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization
Key Information Extraction (KIE) underpins the understanding of visual documents (e.g., receipts and contracts) by extracting precise semantic content and accurately capturing spatial structure. Yet existing multimodal l…
Key Information ExtractionPaddleOCR 3.0 Technical Report
This technical report introduces PaddleOCR 3.0, an Apache-licensed open-source toolkit for OCR and document parsing. To address the growing demand for document understanding in the era of large language models, PaddleOCR…
document understandingKey Information ExtractionOptical Character Recognition (OCR)Class-Agnostic Region-of-Interest Matching in Document Images
Document understanding and analysis have received a lot of attention due to their widespread application. However, existing document analysis solutions, such as document layout analysis and key information extraction, ar…
Document Layout Analysisdocument understandingKey Information Extraction