From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
The Controllable Image Captioning Agent (CapAgent) is an innovative system designed to bridge the gap between user simplicity and professional-level outputs in image captioning tasks. CapAgent automatically transforms user-provided simple instructions into detailed, professional instructions, enabling precise and context-aware caption generation. By leveraging multimodal large language models (MLLMs) and external tools such as object detection tool and search engines, the system ensures that captions adhere to specified guidelines, including sentiment, keywords, focus, and formatting. CapAgent transparently controls each step of the captioning process, and showcases its reasoning and tool usage at every step, fostering user trust and engagement. The project code is available at https://github.com/xin-ran-w/CapAgent.
Code (1)
Tasks
Caption Generationcontrollable image captioningImage Captioningobject-detectionObject DetectionSimilar Papers 제목 키워드 기반
Learning Combinatorial Prompts for Universal Controllable Image Captioning
Controllable Image Captioning (CIC) -- generating natural language descriptions about images under the guidance of given control signals -- is one of the most promising directions towards next-generation captioning syste…
controllable image captioningImage CaptioningLanguage ModelingLanguage Modelling+1Length-Controllable Image Captioning
The last decade has witnessed remarkable progress in the image captioning task; however, most existing methods cannot control their captions, \emph{e.g.}, choosing to describe the image either roughly or in detail. In th…
controllable image captioningDecoderDiversityImage CaptioningCineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning
Cinematographic captioning aims to describe how a video is filmed using professional film-language concepts such as camera movement, shot size, depth of field, composition, and shooting angle. This capability is importan…
Reinforcement LearningVideo CaptioningVideo GenerationControllable Image Captioning via Prompting
Despite the remarkable progress of image captioning, existing captioners typically lack the controllable capability to generate desired image captions, e.g., describing the image in a rough or detailed manner, in a factu…
controllable image captioningImage CaptioningPrompt EngineeringPrompt LearningLanguage-Driven Region Pointer Advancement for Controllable Image Captioning
Controllable Image Captioning is a recent sub-field in the multi-modal task of Image Captioning wherein constraints are placed on which regions in an image should be described in the generated natural language caption. T…
controllable image captioningImage CaptioningSentence