paper-with-me

Papers

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures

2026-07-29 · Fahad Ahmed, Sören Auer, Jennifer D'Souza arxiv

Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset, accompanied by a community-driven competition, raises the bar over prior scientific competitions by curating a comprehensive, expert-annotated dataset across four end-to-end complementary tasks. The competition attracted 68 active participants and 1,263 public/private submissions from 9th January 2026 to 8th April 2026. Our results show that state-of-the-art multimodal models perform well on classification and summarization tasks but struggle with data extraction and scientific reasoning, particularly in visual question-answering. These findings reveal key limitations and highlight challenges and opportunities for improving domain-aware multimodal AI systems. Overall, the Sci-ImageMiner benchmark and competition establish a rigorous platform for advancing research in scientific figure comprehension and reasoning and demonstrate the potential of state-of-the-art approaches for a challenging and complex research area.

📄 PDF Abstract BibTeX arXiv:2607.26848

Code (0)

등록된 구현이 없습니다.

Tasks

Information Extraction

Similar Papers 제목 키워드 기반

System Description of CITlab's Recognition & Retrieval Engine for ICDAR2017 Competition on Information Extraction in Historical Handwritten Records

2018-04-26 · Strauß Tobias, Weidemann Max, Michael Johannes, Leifert Gundram 외

We present a recognition and retrieval system for the ICDAR2017 Competition on Information Extraction in Historical Handwritten Records which successfully infers person names and other data from marriage records. The sys…

Retrieval

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

2026-08-14 · Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova 외 arxiv

Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competit…

Visual Question AnsweringInformation Extraction

ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction

2021-03-18 · Zheng Huang, Kai Chen, Jianhua He, Xiang Bai 외

Scanned receipts OCR and key information extraction (SROIE) represent the processeses of recognizing text from scanned receipts and extracting key texts from them and save the extracted tests to structured documents. SRO…

Key Information ExtractionOptical Character Recognition (OCR)Task 2

ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images

2023-06-05 · Wenwen Yu, Chengquan Zhang, Haoyu Cao, Wei Hua 외

Structured text extraction is one of the most valuable and challenging application directions in the field of Document AI. However, the scenarios of past benchmarks are limited, and the corresponding evaluation protocols…

Document AIEntity Linkingvalid

Baseline Detection in Historical Documents using Convolutional U-Nets

2018-10-22 · Michael Fink, Thomas Layer, Georg Mackenbrock, Michael Sprinzl

Baseline detection is still a challenging task for heterogeneous collections of historical documents. We present a novel approach to baseline extraction in such settings, turning out the winning entry to the ICDAR 2017 C…