paper-with-me

Papers

Grounding Physical Object and Event Concepts Through Dynamic Visual Reasoning

2021-01-01 · ICLR 2021 1 · Zhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee Kenneth Wong, Joshua B. Tenenbaum, Chuang Gan

We study the problem of dynamic visual reasoning on raw videos. This is a challenging problem; currently, state-of-the-art models often require dense supervision on physical object properties and events from simulation, which are impractical to obtain in real life. In this paper, we present the Dynamic Concept Learner (DCL), a unified framework that grounds physical objects and events from video and language. DCL first adopts a trajectory extractor to track each object over time and to represent it as a latent, object-centric feature vector. Building upon this object-centric representation, DCL learns to approximate the dynamic interaction among objects using graph networks. DCL further incorporates a semantic parser to parse question into semantic programs and, finally, a program executor to run the program to answer the question, levering the learned dynamics model. After training, DCL can detect and associate objects across the frames, ground visual properties and physical events, understand the causal relationship between events, make future and counterfactual predictions, and leverage these extracted presentations for answering queries. DCL achieves state-of-the-art performance on CLEVRER, a challenging causal video reasoning dataset, even without using ground-truth attributes and collision labels from simulations for training. We further test DCL on a newly proposed video-retrieval and event localization dataset derived from CLEVRER, showing its strong generalization capacity.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualObjectRetrievalVideo RetrievalVisual Reasoning

Similar Papers 제목 키워드 기반

Grounding Physical Concepts of Objects and Events Through Dynamic Visual Reasoning

2021-03-30 · Zhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee Kenneth Wong 외

We study the problem of dynamic visual reasoning on raw videos. This is a challenging problem; currently, state-of-the-art models often require dense supervision on physical object properties and events from simulation, …

counterfactualObjectRetrievalVideo Retrieval+1

VideoGEM: Training-free Action Grounding in Videos

2025-03-26 · CVPR 2025 1 · Felix Vogel, Walid Bousselham, Anna Kukleva, Nina Shvetsova 외

Vision-language foundation models have shown impressive capabilities across various zero-shot tasks, including training-free localization and grounding, primarily focusing on localizing objects in images. However, levera…

Video Grounding

Dual-Channel Grounded World Modeling (DCGWM): Structural Prevention of Objective Interference Collapse via Heterogeneous External Grounding with Inward-Only Gradient Flow

2026-06-17 · Akshay Hazare arxiv

Joint Embedding Predictive Architectures (JEPAs) are a leading approach to world model representation learning. We identify a failure mode in JEPA-based world models grounded against two qualitatively distinct external s…

Representation Learning

Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts

2025-03-30 · Jianhua Sun, Jiude Wei, YuXuan Li, Cewu Lu

We human rely on a wide range of commonsense knowledge to interact with an extensive number and categories of objects in the physical world. Likewise, such commonsense knowledge is also crucial for robots to successfully…

Object

Executable Analytic Concepts as the Missing Link Between VLM Insight and Precise Manipulation

2025-10-09 · Mingyang Sun, Jiude Wei, Qichen He, Donglin Wang 외 arxiv

Enabling robots to perform precise and generalized manipulation in unstructured environments remains a fundamental challenge in embodied AI. While Vision-Language Models (VLMs) have demonstrated remarkable capabilities i…

Zero-shot Generalization