paper-with-me

홈 › Papers

Context and Attribute Grounded Dense Captioning

2019-04-02 · CVPR 2019 6 · Guojun Yin, Lu Sheng, Bin Liu, Nenghai Yu, Xiaogang Wang, Jing Shao

Dense captioning aims at simultaneously localizing semantic regions and describing these regions-of-interest (ROIs) with short phrases or sentences in natural language. Previous studies have shown remarkable progresses, but they are often vulnerable to the aperture problem that a caption generated by the features inside one ROI lacks contextual coherence with its surrounding context in the input image. In this work, we investigate contextual reasoning based on multi-scale message propagations from the neighboring contents to the target ROIs. To this end, we design a novel end-to-end context and attribute grounded dense captioning framework consisting of 1) a contextual visual mining module and 2) a multi-level attribute grounded description generation module. Knowing that captions often co-occur with the linguistic attributes (such as who, what and where), we also incorporate an auxiliary supervision from hierarchical linguistic attributes to augment the distinctiveness of the learned captions. Extensive experiments and ablation studies on Visual Genome dataset demonstrate the superiority of the proposed model in comparison to state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:1904.01410

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeDense Captioning

Similar Papers 제목 키워드 기반

See It All: Contextualized Late Aggregation for 3D Dense Captioning

2024-08-14 · Minjung Kim, Hyung Suk Lim, Seung Hwan Kim, Soonyoung Lee 외

3D dense captioning is a task to localize objects in a 3D scene and generate descriptive sentences for each object. Recent approaches in 3D dense captioning have adopted transformer encoder-decoder frameworks from object…

3D dense captioningAllAttributeCaption Generation+6

ComiCap: A VLMs pipeline for dense captioning of Comic Panels

2024-09-24 · Emanuele Vivoli, Niccolò Biondi, Marco Bertini, Dimosthenis Karatzas

The comic domain is rapidly advancing with the development of single- and multi-page analysis and synthesis models. Recent benchmarks and datasets have been introduced to support and assess models' capabilities in tasks …

AttributeDense CaptioningSpeaker Identification

Bi-directional Contextual Attention for 3D Dense Captioning

2024-08-13 · Minjung Kim, Hyung Suk Lim, Soonyoung Lee, Bumsoo Kim 외

3D dense captioning is a task involving the localization of objects and the generation of descriptions for each object in a 3D scene. Recent approaches have attempted to incorporate contextual information by modeling rel…

3D dense captioningAttributeCaption GenerationDense Captioning+1

Activitynet 2019 Task 3: Exploring Contexts for Dense Captioning Events in Videos

2019-07-11 · Shizhe Chen, Yuqing Song, Yida Zhao, Qin Jin 외

Contextual reasoning is essential to understand events in long untrimmed videos. In this work, we systematically explore different captioning models with various contexts for the dense-captioning events in video task, wh…

Dense CaptioningDense Video CaptioningDiversityVideo Captioning

HallE-Control: Controlling Object Hallucination in Large Multimodal Models

2023-10-03 · Bohan Zhai, Shijia Yang, Chenfeng Xu, Sheng Shen 외

Current Large Multimodal Models (LMMs) achieve remarkable progress, yet there remains significant uncertainty regarding their ability to accurately apprehend visual details, that is, in performing detailed captioning. To…

AttributeDecoderHallucinationObject+2