paper-with-me

홈 › Papers

Learning Descriptive Image Captioning via Semipermeable Maximum Likelihood Estimation

2023-09-21 · NeurIPS 2023 11

Image captioning aims to describe visual content in natural language. As 'a picture is worth a thousand words', there could be various correct descriptions for an image. However, with maximum likelihood estimation as the training objective, the captioning model is penalized whenever its prediction mismatches with the label. For instance, when the model predicts a word expressing richer semantics than the label, it will be penalized and optimized to prefer more concise expressions, referred to as *conciseness optimization*. In contrast, predictions that are more concise than labels lead to *richness optimization*. Such conflicting optimization directions could eventually result in the model generating general descriptions. In this work, we introduce Semipermeable MaxImum Likelihood Estimation (SMILE), which allows richness optimization while blocking conciseness optimization, thus encouraging the model to generate longer captions with more details. Extensive experiments on two mainstream image captioning datasets MSCOCO and Flickr30K demonstrate that SMILE significantly enhances the descriptiveness of generated captions. We further provide in-depth investigations to facilitate a better understanding of how SMILE works.Submission Number: 10080

📄 PDF Abstract BibTeX

Code (1)

yuezih/smile 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Enhancing Descriptive Image Captioning with Natural Language Inference

2021-08-01 · ACL 2021 5 · Zhan Shi, Hui Liu, Xiaodan Zhu

Generating \textit{descriptive} sentences that convey non-trivial, detailed, and salient information about images is an important goal of image captioning. In this paper we propose a novel approach to encourage captionin…

DescriptiveImage CaptioningNatural Language Inference

Towards Unique and Informative Captioning of Images

2020-09-08 · ECCV 2020 8 · Zeyu Wang, Berthy Feng, Karthik Narasimhan, Olga Russakovsky

Despite considerable progress, state of the art image captioning models produce generic captions, leaving out important image details. Furthermore, these systems may even misrepresent the image in order to produce a simp…

DiversityImage CaptioningRe-Ranking

Image Captioners Sometimes Tell More Than Images They See

2023-05-04 · Honori Udo, Takafumi Koshinaka

Image captioning, a.k.a. "image-to-text," which generates descriptive text from given images, has been rapidly developing throughout the era of deep learning. To what extent is the information in the original image prese…

DescriptiveImage Captioningimage-classificationImage Classification+1

CIC: A Framework for Culturally-Aware Image Captioning

2024-02-08 · Youngsik Yun, Jihie Kim

Image Captioning generates descriptive sentences from images using Vision-Language Pre-trained models (VLPs) such as BLIP, which has improved greatly. However, current methods lack the generation of detailed descriptive …

DescriptiveImage CaptioningQuestion AnsweringVisual Question Answering+1

Controlling Length in Image Captioning

2020-05-29 · Ruotian Luo, Greg Shakhnarovich

We develop and evaluate captioning models that allow control of caption length. Our models can leverage this control to generate captions of different style and descriptiveness.

Image Captioning