paper-with-me

Papers

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity

2024-11-23 · CVPR 2025 1 · Hang Hua, Qing Liu, Lingzhi Zhang, Jing Shi, Zhifei Zhang, Yilin Wang, Jianming Zhang, Jiebo Luo

The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate reasoning across various applications, including image and video captioning, visual question answering, and cross-modal retrieval. Despite their superior capabilities, VLMs struggle with fine-grained image regional composition information perception. Specifically, they have difficulty accurately aligning the segmentation masks with the corresponding semantics and precisely describing the compositional aspects of the referred regions. However, compositionality - the ability to understand and generate novel combinations of known visual and textual components - is critical for facilitating coherent reasoning and understanding across modalities by VLMs. To address this issue, we propose FINECAPTION, a novel VLM that can recognize arbitrary masks as referential inputs and process high-resolution images for compositional image captioning at different granularity levels. To support this endeavor, we introduce COMPOSITIONCAP, a new dataset for multi-grained region compositional image captioning, which introduces the task of compositional attribute-aware regional image captioning. Empirical results demonstrate the effectiveness of our proposed model compared to other state-of-the-art VLMs. Additionally, we analyze the capabilities of current VLMs in recognizing various visual prompts for compositional region image captioning, highlighting areas for improvement in VLM design and training.

📄 PDF Abstract BibTeX arXiv:2411.15411

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeCross-Modal RetrievalImage CaptioningQuestion AnsweringVideo CaptioningVisual Question Answering

Similar Papers 제목 키워드 기반

A Neural Compositional Paradigm for Image Captioning

2018-10-23 · NeurIPS 2018 12 · Bo Dai, Sanja Fidler, Dahua Lin

Mainstream captioning models often follow a sequential structure to generate captions, leading to issues such as introduction of irrelevant semantics, lack of diversity in the generated captions, and inadequate generaliz…

DiversityImage Captioning

A Neural Compositional Paradigm for Image Captioning

2018-09-24 · Anonymous

Mainstream captioning models often follow a sequential structure to generate cap- tions, leading to issues such as introduction of irrelevant semantics, lack of diversity in the generated captions, and inadequate general…

DiversityImage Captioning

The Role of Syntactic Planning in Compositional Image Captioning

2021-01-28 · EACL 2021 2 · Emanuele Bugliarello, Desmond Elliott

Image captioning has focused on generalizing to images drawn from the same distribution as the training set, and not to the more challenging problem of generalizing to different distributions of images. Recently, Nikolau…

Image Captioning

Compositional Generalization in Image Captioning

2019-09-10 · CONLL 2019 11 · Mitja Nikolaus, Mostafa Abdou, Matthew Lamm, Rahul Aralikatte 외

Image captioning models are usually evaluated on their ability to describe a held-out set of images, not on their ability to generalize to unseen concepts. We study the problem of compositional generalization, which meas…

Caption GenerationImage CaptioningSentence

Image Captioning with Compositional Neural Module Networks

2020-07-10 · Junjiao Tian, Jean Oh

In image captioning where fluency is an important factor in evaluation, e.g., $n$-gram metrics, sequential models are commonly used; however, sequential models generally result in overgeneralized expressions that lack th…

Image CaptioningQuestion AnsweringSentenceVisual Question Answering+1