paper-with-me

홈 › Papers

Weakly-Supervised Generation and Grounding of Visual Descriptions With Conditional Generative Models

2022-01-01 · CVPR 2022 1 · Effrosyni Mavroudi, René Vidal

Given weak supervision from image- or video-caption pairs, we address the problem of grounding (localizing) each object word of a ground-truth or generated sentence describing a visual input. Recent weakly-supervised approaches leverage region proposals and ground words based on the region attention coefficients of captioning models. To predict each next word in the sentence they attend over regions using a summary of the previous words as a query, and then ground the word by selecting the most attended regions. However, this leads to sub-optimal grounding, since attention coefficients are computed without taking into account the word that needs to be localized. To address this shortcoming, we propose a novel Grounded Visual Description Conditional Variational Autoencoder (GVD-CVAE) and leverage its latent variables for grounding. In particular, we introduce a discrete random variable that models each word-to-region alignment, and learn its approximate posterior distribution given the full sentence. Experiments on challenging image and video datasets (Flickr30k Entities, YouCook2, ActivityNet Entities) validate the effectiveness of our conditional generative model, showing that it can substantially outperform soft-attention-based baselines in grounding.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment

2025-05-03 · Xiaoqi Li, Jiaming Liu, Nuowei Han, Liang Heng 외

The 3D weakly-supervised visual grounding task aims to localize oriented 3D boxes in point clouds based on natural language descriptions without requiring annotations to guide model learning. This setting presents two pr…

SentenceVisual Grounding

AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding

2025-08-05 · Yidan Wang, Chenyi Zhuang, Wutao Liu, Pan Gao 외 arxiv

Weakly supervised visual grounding (VG) aims to locate objects in images based on text descriptions. Despite significant progress, existing methods lack strong cross-modal reasoning to distinguish subtle semantic differe…

Contrastive LearningVisual Grounding

Modularized Textual Grounding for Counterfactual Resilience

2019-04-07 · CVPR 2019 6 · Zhiyuan Fang, Shu Kong, Charless Fowlkes, Yezhou Yang

Computer Vision applications often require a textual grounding module with precision, interpretability, and resilience to counterfactual inputs/queries. To achieve high grounding precision, current textual grounding meth…

AttributecounterfactualNatural Language Visual GroundingPhrase Grounding+2

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

2024-12-05 · CVPR 2025 1 · Rong Li, Shijie Li, Lingdong Kong, Xulei Yang 외

3D Visual Grounding (3DVG) aims to locate objects in 3D scenes based on textual descriptions, essential for applications like augmented reality and robotics. Traditional 3DVG approaches rely on annotated 3D datasets and …

3D visual groundingObject LocalizationVisual Grounding

Cycle-Consistency Learning for Captioning and Grounding

2023-12-23 · Ning Wang, Jiajun Deng, Mingbo Jia

We present that visual grounding and image captioning, which perform as two mutually inverse processes, can be bridged together for collaborative training by careful designs. By consolidating this idea, we introduce CyCo…

Image CaptioningVisual Grounding