paper-with-me

홈 › Papers

From Simple to Professional: A Combinatorial Controllable Image Captioning Agent

2024-12-15 · Xinran Wang, Muxi Diao, Baoteng Li, Haiwen Zhang, Kongming Liang, Zhanyu Ma

The Controllable Image Captioning Agent (CapAgent) is an innovative system designed to bridge the gap between user simplicity and professional-level outputs in image captioning tasks. CapAgent automatically transforms user-provided simple instructions into detailed, professional instructions, enabling precise and context-aware caption generation. By leveraging multimodal large language models (MLLMs) and external tools such as object detection tool and search engines, the system ensures that captions adhere to specified guidelines, including sentiment, keywords, focus, and formatting. CapAgent transparently controls each step of the captioning process, and showcases its reasoning and tool usage at every step, fostering user trust and engagement. The project code is available at https://github.com/xin-ran-w/CapAgent.

📄 PDF Abstract BibTeX arXiv:2412.11025

Code (1)

xin-ran-w/capagent 공식 구현

Tasks

Caption Generationcontrollable image captioningImage Captioningobject-detectionObject Detection

Similar Papers 제목 키워드 기반

Learning Combinatorial Prompts for Universal Controllable Image Captioning

2023-03-11 · Zhen Wang, Jun Xiao, Yueting Zhuang, Fei Gao 외

Controllable Image Captioning (CIC) -- generating natural language descriptions about images under the guidance of given control signals -- is one of the most promising directions towards next-generation captioning syste…

controllable image captioningImage CaptioningLanguage ModelingLanguage Modelling+1

Length-Controllable Image Captioning

2020-07-19 · ECCV 2020 8 · Chaorui Deng, Ning Ding, Mingkui Tan, Qi Wu

The last decade has witnessed remarkable progress in the image captioning task; however, most existing methods cannot control their captions, \emph{e.g.}, choosing to describe the image either roughly or in detail. In th…

controllable image captioningDecoderDiversityImage Captioning

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning

2026-06-23 · Xinyu Mao, Yuhui Zeng, Xiaokun Liu, Wenyu Qin 외 arxiv

Cinematographic captioning aims to describe how a video is filmed using professional film-language concepts such as camera movement, shot size, depth of field, composition, and shooting angle. This capability is importan…

Reinforcement LearningVideo CaptioningVideo Generation

Controllable Image Captioning via Prompting

2022-12-04 · Ning Wang, Jiahao Xie, Jihao Wu, Mingbo Jia 외

Despite the remarkable progress of image captioning, existing captioners typically lack the controllable capability to generate desired image captions, e.g., describing the image in a rough or detailed manner, in a factu…

controllable image captioningImage CaptioningPrompt EngineeringPrompt Learning

Language-Driven Region Pointer Advancement for Controllable Image Captioning

2020-11-30 · COLING 2020 8 · Annika Lindh, Robert J. Ross, John D. Kelleher

Controllable Image Captioning is a recent sub-field in the multi-modal task of Image Captioning wherein constraints are placed on which regions in an image should be described in the generated natural language caption. T…

controllable image captioningImage CaptioningSentence