paper-with-me

Papers

SciCap+: A Knowledge Augmented Dataset to Study the Challenges of Scientific Figure Captioning

2023-06-06 · Zhishen Yang, Raj Dabre, Hideki Tanaka, Naoaki Okazaki

In scholarly documents, figures provide a straightforward way of communicating scientific findings to readers. Automating figure caption generation helps move model understandings of scientific documents beyond text and will help authors write informative captions that facilitate communicating scientific findings. Unlike previous studies, we reframe scientific figure captioning as a knowledge-augmented image captioning task that models need to utilize knowledge embedded across modalities for caption generation. To this end, we extended the large-scale SciCap dataset~\cite{hsu-etal-2021-scicap-generating} to SciCap+ which includes mention-paragraphs (paragraphs mentioning figures) and OCR tokens. Then, we conduct experiments with the M4C-Captioner (a multimodal transformer-based model with a pointer network) as a baseline for our study. Our results indicate that mention-paragraphs serves as additional context knowledge, which significantly boosts the automatic standard image caption evaluation scores compared to the figure-only baselines. Human evaluations further reveal the challenges of generating figure captions that are informative to readers. The code and SciCap+ dataset will be publicly available at https://github.com/ZhishenYang/scientific_figure_captioning_dataset

📄 PDF Abstract BibTeX arXiv:2306.03491

Code (1)

zhishenyang/scientific_figure_captioning_dataset 공식 구현

Tasks

Caption GenerationImage CaptioningOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

SciCapenter: Supporting Caption Composition for Scientific Figures with Machine-Generated Captions and Ratings

2024-03-26 · Ting-Yao Hsu, Chieh-Yang Huang, Shih-Hong Huang, Ryan Rossi 외

Crafting effective captions for figures is important. Readers heavily depend on these captions to grasp the figure's message. However, despite a well-developed set of AI technologies for figures and captions, these have …

Optical Character Recognition (OCR)

SciCap: Generating Captions for Scientific Figures

2021-10-22 · Findings (EMNLP) 2021 11 · Ting-Yao Hsu, C. Lee Giles, Ting-Hao 'Kenneth' Huang

Researchers use figures to communicate rich, complex information in scientific papers. The captions of these figures are critical to conveying effective messages. However, low-quality figure captions commonly occur in sc…

ArticlesImage CaptioningText Normalization

Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SCICAP Challenge 2023

2025-01-31 · Ting-Yao E. Hsu, Yi-Li Hsu, Shaurya Rohatgi, Chieh-Yang Huang 외

Since the SCICAP datasets launch in 2021, the research community has made significant progress in generating captions for scientific figures in scholarly articles. In 2023, the first SCICAP Challenge took place, inviting…

ArticlesCaption GenerationText Generation

Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning

2025-12-25 · Ting-Hao 'Kenneth' Huang, Ryan A. Rossi, Sungchul Kim, Tong Yu 외 arxiv

Between 2021 and 2025, the SciCap project grew from a small seed-funded idea at The Pennsylvania State University (Penn State) into one of the central efforts shaping the scientific figure-captioning landscape. Supported…

Proposal Report for the 2nd SciCAP Competition 2024

2024-07-02 · Pengpeng Li, Tingmin Li, Jingyuan Wang, Boyuan Wang 외

In this paper, we propose a method for document summarization using auxiliary information. This approach effectively summarizes descriptions related to specific images, tables, and appendices within lengthy texts. Our ex…

Document SummarizationOptical Character Recognition (OCR)Text Generation