Key Information Extraction
6개 벤치마크 · 논문 92편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
LayoutLM: Pre-training of Text and Layout for Document Image Understanding
LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding
LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking
Papers
DocClaw: A Unified Agentic System for Intelligent Document Processing
Intelligent document processing (IDP) encompasses a broad range of tasks, including optical character recognition (OCR), document question answering (DocQA), and key information extraction (KIE). Despite their distinct o…
Key Information ExtractionQuestion AnsweringBeyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs
While the privacy risks of multimodal large language models (MLLMs) have drawn significant attention, the unique vulnerabilities of domain-specific MLLMs remain largely underexplored. Focusing on document understanding M…
Key Information ExtractionEnhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis
Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-server Large Multimodal Models (LMMs), wh…
Key Information ExtractionBaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension
Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing adoption in real-world, human-centric applications. However, this adoption is limit…
Key Information ExtractionDocument Layout AnalysisCC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing
Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capability in document literacy. However, their effectiveness in real-worl…
Key Information ExtractionQuestion AnsweringJaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
Japanese scene text poses challenges that multilingual benchmarks often fail to capture, including mixed scripts, frequent vertical writing, and a character inventory far larger than the Latin alphabet. Although Japanese…
Key Information ExtractionVisual Question Answering