paper-with-me

Papers

Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SCICAP Challenge 2023

2025-01-31 · Ting-Yao E. Hsu, Yi-Li Hsu, Shaurya Rohatgi, Chieh-Yang Huang, Ho Yin Sam Ng, Ryan Rossi, Sungchul Kim, Tong Yu, Lun-Wei Ku, C. Lee Giles, Ting-Hao K. Huang

Since the SCICAP datasets launch in 2021, the research community has made significant progress in generating captions for scientific figures in scholarly articles. In 2023, the first SCICAP Challenge took place, inviting global teams to use an expanded SCICAP dataset to develop models for captioning diverse figure types across various academic fields. At the same time, text generation models advanced quickly, with many powerful pre-trained large multimodal models (LMMs) emerging that showed impressive capabilities in various vision-and-language tasks. This paper presents an overview of the first SCICAP Challenge and details the performance of various models on its data, capturing a snapshot of the fields state. We found that professional editors overwhelmingly preferred figure captions generated by GPT-4V over those from all other models and even the original captions written by authors. Following this key finding, we conducted detailed analyses to answer this question: Have advanced LMMs solved the task of generating captions for scientific figures?

📄 PDF Abstract BibTeX arXiv:2501.19353

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesCaption GenerationText Generation

Similar Papers 제목 키워드 기반

Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer

2025-09-30 · Jaeyoung Kim, Jongho Lee, Hongjun Choi, Sion Jang arxiv

We study personalized figure caption generation using author profile data from scientific papers. Our experiments demonstrate that rich author profile data, combined with relevant metadata, can significantly improve the …

Style Transfer

SciCap+: A Knowledge Augmented Dataset to Study the Challenges of Scientific Figure Captioning

2023-06-06 · Zhishen Yang, Raj Dabre, Hideki Tanaka, Naoaki Okazaki

In scholarly documents, figures provide a straightforward way of communicating scientific findings to readers. Automating figure caption generation helps move model understandings of scientific documents beyond text and …

Caption GenerationImage CaptioningOptical Character Recognition (OCR)

Multi-LLM Collaborative Caption Generation in Scientific Documents

2025-01-05 · Jaeyoung Kim, Jongho Lee, Hong-Jun Choi, Ting-Yao Hsu 외

Scientific figure captioning is a complex task that requires generating contextually appropriate descriptions of visual content. However, existing methods often fall short by utilizing incomplete information, treating th…

Caption GenerationImage to textText Summarization

MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding

2024-07-06 · Zekun Li, Xianjun Yang, Kyuri Choi, Wanrong Zhu 외

The rapid development of Multimodal Large Language Models (MLLMs) is making AI-driven scientific assistants increasingly feasible, with interpreting scientific figures being a crucial task. However, existing datasets and…

ArticlesInstruction FollowingMultiple-choicevisual instruction following

S1-MMAlign: A Large-Scale, Multi-Disciplinary Dataset for Scientific Figure-Text Understanding

2026-01-01 · He Wang, Longteng Guo, Pengkang Huo, Xuanxu Lin 외 arxiv

Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. We pre…