paper-with-me

Papers

Concadia: Towards Image-Based Text Generation with a Purpose

2021-04-16 · Elisa Kreiss, Fei Fang, Noah D. Goodman, Christopher Potts

Current deep learning models often achieve excellent results on benchmark image-to-text datasets but fail to generate texts that are useful in practice. We argue that to close this gap, it is vital to distinguish descriptions from captions based on their distinct communicative roles. Descriptions focus on visual features and are meant to replace an image (often to increase accessibility), whereas captions appear alongside an image to supply additional information. To motivate this distinction and help people put it into practice, we introduce the publicly available Wikipedia-based dataset Concadia consisting of 96,918 images with corresponding English-language descriptions, captions, and surrounding context. Using insights from Concadia, models trained on it, and a preregistered human-subjects experiment with human- and model-generated texts, we characterize the commonalities and differences between descriptions and captions. In addition, we show that, for generating both descriptions and captions, it is useful to augment image-to-text models with representations of the textual context in which the image appeared.

📄 PDF Abstract BibTeX arXiv:2104.08376

Code (1)

elisakreiss/concadia 공식 구현 pytorch

Tasks

Image CaptioningImage to textText Generation

Similar Papers 제목 키워드 기반

VLIS: Unimodal Language Models Guide Multimodal Language Generation

2023-10-15 · Jiwan Chung, Youngjae Yu

Multimodal language generation, which leverages the synergy of language and vision, is a rapidly expanding field. However, existing vision-language models face challenges in tasks that require complex linguistic understa…

Caption GenerationExplanation GenerationImage Paragraph CaptioningLanguage Modeling+4

Updating CLIP to Prefer Descriptions Over Captions

2024-06-12 · Amir Zur, Elisa Kreiss, Karel D'Oosterlinck, Christopher Potts 외

Although CLIPScore is a powerful generic metric that captures the similarity between a text and an image, it fails to distinguish between a caption that is meant to complement the information in an image and a descriptio…

parameter-efficient fine-tuning

BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models

2023-12-05 · CVPR 2024 1 · Fengyuan Shi, Jiaxi Gu, Hang Xu, Songcen Xu 외

Diffusion models have made tremendous progress in text-driven image and video generation. Now text-to-image foundation models are widely applied to various downstream image synthesis tasks, such as controllable image gen…

Image GenerationModel SelectionVideo EditingVideo Generation+1

DuoGen: Towards General Purpose Interleaved Multimodal Generation

2026-01-31 · Min Shi, Xiaohui Zeng, Jiannan Huang, Yin Cui 외 arxiv

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts for reasoning. However, the quality of ex…

multimodal generationVideo GenerationImage Editing

IR image databases generation under target intrinsic thermal variability constraints

2024-11-12 · Jerome Gilles, Stephane Landeau, Tristan Dagobert, Philippe Chevalier 외

This paper deals with the problem of infrared image database generation for ATR assessment purposes. Huge databases are required to have quantitative and objective performance evaluations. We propose a method which super…