paper-with-me

홈 › Papers

Controllable Image Captioning via Prompting

2022-12-04 · Ning Wang, Jiahao Xie, Jihao Wu, Mingbo Jia, Linlin Li

Despite the remarkable progress of image captioning, existing captioners typically lack the controllable capability to generate desired image captions, e.g., describing the image in a rough or detailed manner, in a factual or emotional view, etc. In this paper, we show that a unified model is qualified to perform well in diverse domains and freely switch among multiple styles. Such a controllable capability is achieved by embedding the prompt learning into the image captioning framework. To be specific, we design a set of prompts to fine-tune the pre-trained image captioner. These prompts allow the model to absorb stylized data from different domains for joint training, without performance degradation in each domain. Furthermore, we optimize the prompts with learnable vectors in the continuous word embedding space, avoiding the heuristic prompt engineering and meanwhile exhibiting superior performance. In the inference stage, our model is able to generate desired stylized captions by choosing the corresponding prompts. Extensive experiments verify the controllable capability of the proposed method. Notably, we achieve outstanding performance on two diverse image captioning benchmarks including COCO Karpathy split and TextCaps using a unified model.

📄 PDF Abstract BibTeX arXiv:2212.01803

Code (0)

등록된 구현이 없습니다.

Tasks

controllable image captioningImage CaptioningPrompt EngineeringPrompt Learning

Similar Papers 제목 키워드 기반

Controllable Contextualized Image Captioning: Directing the Visual Narrative through User-Defined Highlights

2024-07-16 · Shunqi Mao, Chaoyi Zhang, Hang Su, Hwanjun Song 외

Contextualized Image Captioning (CIC) evolves traditional image captioning into a more complex domain, necessitating the ability for multimodal reasoning. It aims to generate image captions given specific contextual info…

Image CaptioningMultimodal Reasoning

Language-Driven Region Pointer Advancement for Controllable Image Captioning

2020-11-30 · COLING 2020 8 · Annika Lindh, Robert J. Ross, John D. Kelleher

Controllable Image Captioning is a recent sub-field in the multi-modal task of Image Captioning wherein constraints are placed on which regions in an image should be described in the generated natural language caption. T…

controllable image captioningImage CaptioningSentence

Show, Control and Tell: A Framework for Generating Controllable and Grounded Captions

2018-11-26 · CVPR 2019 6 · Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

Current captioning approaches can describe images using black-box architectures whose behavior is hardly controllable and explainable from the exterior. As an image can be described in infinite ways depending on the goal…

controllable image captioningDiversityImage Captioning

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent

2024-12-15 · Xinran Wang, Muxi Diao, Baoteng Li, Haiwen Zhang 외

The Controllable Image Captioning Agent (CapAgent) is an innovative system designed to bridge the gap between user simplicity and professional-level outputs in image captioning tasks. CapAgent automatically transforms us…

Caption Generationcontrollable image captioningImage Captioningobject-detection+1

Length-Controllable Image Captioning

2020-07-19 · ECCV 2020 8 · Chaorui Deng, Ning Ding, Mingkui Tan, Qi Wu

The last decade has witnessed remarkable progress in the image captioning task; however, most existing methods cannot control their captions, \emph{e.g.}, choosing to describe the image either roughly or in detail. In th…

controllable image captioningDecoderDiversityImage Captioning