paper-with-me

Key Information Extraction

6개 벤치마크 · 논문 92편 · 이 태스크의 논문 보기 →

Benchmarks

CORD

결과 9개

SROIE

결과 5개

Kleister NDA

결과 3개

EPHOIE

결과 1개

ETD500

결과 1개

SIMARA

결과 1개

Most implemented

Papers

DocClaw: A Unified Agentic System for Intelligent Document Processing

2026-08-19 · Siqi Xiang, Zhipeng Xu, Yufei Liu, Junhao Ji 외 arxiv

Intelligent document processing (IDP) encompasses a broad range of tasks, including optical character recognition (OCR), document question answering (DocQA), and key information extraction (KIE). Despite their distinct o…

Key Information ExtractionQuestion Answering

Beyond Visual Evidence: Revealing and Mitigating Relational Privacy Leakage in Document MLLMs

2026-08-13 · Beining Xu, Hairui Wang, Jiaxin Wang, Changsheng Chen 외 arxiv

While the privacy risks of multimodal large language models (MLLMs) have drawn significant attention, the unique vulnerabilities of domain-specific MLLMs remain largely underexplored. Focusing on document understanding M…

Key Information Extraction

Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis

2026-07-06 · Zhipeng Xu, Zulong Chen, Qing Liu, Junhao Ji 외 arxiv

Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-server Large Multimodal Models (LMMs), wh…

Key Information Extraction

BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension

2026-07-06 · Abu Tyeb Azad, Ishita Sur Apan, Fahim Ahmed, Sumaiya Karim Katha 외 arxiv

Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing adoption in real-world, human-centric applications. However, this adoption is limit…

Key Information ExtractionDocument Layout Analysis

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing

2026-05-05 · Zhipeng Xu, Junhao Ji, Zulong Chen, Zhenghao Liu 외 arxiv

Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capability in document literacy. However, their effectiveness in real-worl…

Key Information ExtractionQuestion Answering

JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding

2026-03-30 · Koki Maeda, Naoaki Okazaki arxiv

Japanese scene text poses challenges that multilingual benchmarks often fail to capture, including mixed scripts, frequent vertical writing, and a character inventory far larger than the Latin alphabet. Although Japanese…

Key Information ExtractionVisual Question Answering

전체 92편 보기 →