paper-with-me

홈 › Papers

Inverse Compositional Learning for Weakly-supervised Relation Grounding

2023-01-01 · ICCV 2023 1 · Huan Li, Ping Wei, Zeyu Ma, Nanning Zheng

Video relation grounding (VRG) is a significant and challenging problem in the domains of cross-modal learning and video understanding. In this study, we introduce a novel approach called inverse compositional learning (ICL) for weakly-supervised video relation grounding. Our approach represents relations at both the holistic and partial levels, formulating VRG as a joint optimization problem that encompasses reasoning at both levels. For holistic-level reasoning, we propose an inverse attention mechanism and a compositional encoder to generate compositional relevance features. Additionally, we introduce an inverse loss to evaluate and learn the relevance between visual features and relation features. At the partial-level reasoning, we introduce a grounding by classification scheme. By leveraging the learned holistic-level features and partial-level features, we train the entire model in an end-to-end manner. We conduct evaluations on two challenging datasets and demonstrate the substantial superiority of our proposed method over state-of-the-art methods. Extensive ablation studies confirm the effectiveness of our approach.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

RelationVideo Understanding

Similar Papers 제목 키워드 기반

Investigating Compositional Challenges in Vision-Language Models for Visual Grounding

2024-01-01 · CVPR 2024 1 · Yunan Zeng, Yan Huang, Jinjin Zhang, Zequn Jie 외

Pre-trained vision-language models (VLMs) have achieved high performance on various downstream tasks which have been widely used for visual grounding tasks in a weakly supervised manner. However despite the performan…

AttributeRelationVisual Grounding

STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding

2025-01-01 · CVPR 2025 1 · Aaryan Garg, Akash Kumar, Yogesh S Rawat

In this work, we study Weakly Supervised Spatio-Temporal Video Grounding (WSTVG), a challenging task of localizing subjects spatio-temporally in videos using only textual queries and no bounding box supervision. Insp…

Action UnderstandingSpatio-Temporal Video GroundingVideo Grounding

Cycle-Consistency Learning for Captioning and Grounding

2023-12-23 · Ning Wang, Jiajun Deng, Mingbo Jia

We present that visual grounding and image captioning, which perform as two mutually inverse processes, can be bridged together for collaborative training by careful designs. By consolidating this idea, we introduce CyCo…

Image CaptioningVisual Grounding

INTRA: Interaction Relationship-aware Weakly Supervised Affordance Grounding

2024-09-10 · Ji Ha Jang, Hoigi Seo, Se Young Chun

Affordance denotes the potential interactions inherent in objects. The perception of affordance can enable intelligent agents to navigate and interact with new environments efficiently. Weakly supervised affordance groun…

Contrastive LearningLanguage ModelingLanguage ModellingNavigate+1

Modularized Textual Grounding for Counterfactual Resilience

2019-04-07 · CVPR 2019 6 · Zhiyuan Fang, Shu Kong, Charless Fowlkes, Yezhou Yang

Computer Vision applications often require a textual grounding module with precision, interpretability, and resilience to counterfactual inputs/queries. To achieve high grounding precision, current textual grounding meth…

AttributecounterfactualNatural Language Visual GroundingPhrase Grounding+2