paper-with-me

홈 › Papers

FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback

2023-07-20 · Ashish Singh, Prateek Agarwal, Zixuan Huang, Arpita Singh, Tong Yu, Sungchul Kim, Victor Bursztyn, Nikos Vlassis, Ryan A. Rossi

Captions are crucial for understanding scientific visualizations and documents. Existing captioning methods for scientific figures rely on figure-caption pairs extracted from documents for training, many of which fall short with respect to metrics like helpfulness, explainability, and visual-descriptiveness [15] leading to generated captions being misaligned with reader preferences. To enable the generation of high-quality figure captions, we introduce FigCaps-HF a new framework for figure-caption generation that can incorporate domain expert feedback in generating captions optimized for reader preferences. Our framework comprises of 1) an automatic method for evaluating quality of figure-caption pairs, 2) a novel reinforcement learning with human feedback (RLHF) method to optimize a generative figure-to-caption model for reader preferences. We demonstrate the effectiveness of our simple learning framework by improving performance over standard fine-tuning across different types of models. In particular, when using BLIP as the base model, our RLHF framework achieves a mean gain of 35.7%, 16.9%, and 9% in ROUGE, BLEU, and Meteor, respectively. Finally, we release a large-scale benchmark dataset with human feedback on figure-caption pairs to enable further evaluation and development of RLHF techniques for this problem.

📄 PDF Abstract BibTeX arXiv:2307.10867

Code (1)

figcapshf/figcapshf 공식 구현 pytorch

Tasks

Caption Generation

Methods 이 논문이 사용한 방법론

BLIP Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based…
BASE 설명 없음

Similar Papers 제목 키워드 기반

SciCap: Generating Captions for Scientific Figures

2021-10-22 · Findings (EMNLP) 2021 11 · Ting-Yao Hsu, C. Lee Giles, Ting-Hao 'Kenneth' Huang

Researchers use figures to communicate rich, complex information in scientific papers. The captions of these figures are critical to conveying effective messages. However, low-quality figure captions commonly occur in sc…

ArticlesImage CaptioningText Normalization

FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures

2026-01-12 · Jifeng Song, Arun Das, Pan Wang, Hui Ji 외 arxiv

Scientific compound figures combine multiple labeled panels into a single image. However, in a PMC-scale crawl of 346,567 compound figures, 16.3% have no caption and 1.8% only have captions shorter than ten words, causin…

Summaries as Captions: Generating Figure Captions for Scientific Documents with Automated Text Summarization

2023-02-23 · Chieh-Yang Huang, Ting-Yao Hsu, Ryan Rossi, Ani Nenkova 외

Good figure captions help paper readers understand complex scientific figures. Unfortunately, even published papers often have poorly written captions. Automatic caption generation could aid paper writers by providing go…

Abstractive Text SummarizationCaption GenerationText Summarization

LineCap: Line Charts for Data Visualization Captioning Models

2022-07-15 · Anita Mahinpei, Zona Kostic, Chris Tanner

Data visualization captions help readers understand the purpose of a visualization and are crucial for individuals with visual impairments. The prevalence of poor figure captions and the successful application of deep le…

Data VisualizationDeep LearningImage Captioning

SciCapenter: Supporting Caption Composition for Scientific Figures with Machine-Generated Captions and Ratings

2024-03-26 · Ting-Yao Hsu, Chieh-Yang Huang, Shih-Hong Huang, Ryan Rossi 외

Crafting effective captions for figures is important. Readers heavily depend on these captions to grasp the figure's message. However, despite a well-developed set of AI technologies for figures and captions, these have …

Optical Character Recognition (OCR)