paper-with-me

홈 › Papers

ComPhy: Compositional Physical Reasoning of Objects and Events from Videos

2022-05-02 · ICLR 2022 4 · Zhenfang Chen, Kexin Yi, Yunzhu Li, Mingyu Ding, Antonio Torralba, Joshua B. Tenenbaum, Chuang Gan

Objects' motions in nature are governed by complex interactions and their properties. While some properties, such as shape and material, can be identified via the object's visual appearances, others like mass and electric charge are not directly visible. The compositionality between the visible and hidden properties poses unique challenges for AI models to reason from the physical world, whereas humans can effortlessly infer them with limited observations. Existing studies on video reasoning mainly focus on visually observable elements such as object appearance, movement, and contact interaction. In this paper, we take an initial step to highlight the importance of inferring the hidden physical properties not directly observable from visual appearances, by introducing the Compositional Physical Reasoning (ComPhy) dataset. For a given set of objects, ComPhy includes few videos of them moving and interacting under different initial conditions. The model is evaluated based on its capability to unravel the compositional hidden properties, such as mass and charge, and use this knowledge to answer a set of questions posted on one of the videos. Evaluation results of several state-of-the-art video reasoning models on ComPhy show unsatisfactory performance as they fail to capture these hidden properties. We further propose an oracle neural-symbolic framework named Compositional Physics Learner (CPL), combining visual perception, physical property learning, dynamic prediction, and symbolic execution into a unified framework. CPL can effectively identify objects' physical properties from their interactions and predict their dynamics to answer questions.

📄 PDF Abstract BibTeX arXiv:2205.01089

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Compositional Physical Reasoning of Objects and Events from Videos

2024-08-02 · Zhenfang Chen, Shilong Dong, Kexin Yi, Yunzhu Li 외

Understanding and reasoning about objects' physical properties in the natural world is a fundamental challenge in artificial intelligence. While some properties like colors and shapes can be directly observed, others, su…

counterfactualQuestion Answering

Intrinsic Physical Concepts Discovery with Object-Centric Predictive Models

2023-03-03 · CVPR 2023 1 · Qu Tang, Xiangyu Zhu, Zhen Lei, Zhaoxiang Zhang

The ability to discover abstract physical concepts and understand how they work in the world through observing lies at the core of human intelligence. The acquisition of this ability is based on compositionally perceivin…

Grounding Physical Concepts of Objects and Events Through Dynamic Visual Reasoning

2021-03-30 · Zhenfang Chen, Jiayuan Mao, Jiajun Wu, Kwan-Yee Kenneth Wong 외

We study the problem of dynamic visual reasoning on raw videos. This is a challenging problem; currently, state-of-the-art models often require dense supervision on physical object properties and events from simulation, …

counterfactualObjectRetrievalVideo Retrieval+1

SPACE: A Simulator for Physical Interactions and Causal Learning in 3D Environments

2021-08-13 · Jiafei Duan, Samson Yu Bai Jian, Cheston Tan

Recent advancements in deep learning, computer vision, and embodied AI have given rise to synthetic causal reasoning video datasets. These datasets facilitate the development of AI algorithms that can reason about physic…

Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering

2024-06-02 · Xingrui Wang, Wufei Ma, Angtian Wang, Shuo Chen 외

For vision-language models (VLMs), understanding the dynamic properties of objects and their interactions within 3D scenes from video is crucial for effective reasoning. In this work, we introduce a video question answer…

counterfactualCounterfactual ReasoningFuture predictionQuestion Answering+1