paper-with-me

Papers

Visual Grounding of Learned Physical Models

2020-04-28 · ICML 2020 1 · Yunzhu Li, Toru Lin, Kexin Yi, Daniel M. Bear, Daniel L. K. Yamins, Jiajun Wu, Joshua B. Tenenbaum, Antonio Torralba

Humans intuitively recognize objects' physical properties and predict their motion, even when the objects are engaged in complicated interactions. The abilities to perform physical reasoning and to adapt to new environments, while intrinsic to humans, remain challenging to state-of-the-art computational models. In this work, we present a neural model that simultaneously reasons about physics and makes future predictions based on visual and dynamics priors. The visual prior predicts a particle-based representation of the system from visual observations. An inference module operates on those particles, predicting and refining estimates of particle locations, object states, and physical parameters, subject to the constraints imposed by the dynamics prior, which we refer to as visual grounding. We demonstrate the effectiveness of our method in environments involving rigid objects, deformable materials, and fluids. Experiments show that our model can infer the physical properties within a few observations, which allows the model to quickly adapt to unseen scenarios and make accurate predictions into the future.

📄 PDF Abstract BibTeX arXiv:2004.13664

Code (1)

yunzhuli/vgpl-dynamics-prior pytorch

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

Neural Material Adaptor for Visual Grounding of Intrinsic Dynamics

2024-10-10 · Junyi Cao, Shanyan Guan, Yanhao Ge, Wei Li 외

While humans effortlessly discern intrinsic dynamics and adapt to new scenarios, modern AI systems often struggle. Current methods for visual grounding of dynamics either use pure neural-network-based simulators (black b…

Visual Grounding

Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding

2023-09-08 · Ozan Unal, Christos Sakaridis, Suman Saha, Luc van Gool

3D visual grounding is the task of localizing the object in a 3D scene which is referred by a description in natural language. With a wide range of applications ranging from autonomous indoor robotics to AR/VR, the task …

3D Instance Segmentation3D visual groundingInstance SegmentationObject+3

Grounding Generated Videos in Feasible Plans via World Models

2026-02-02 · Christos Ziakas, Amir Bar, Alessandra Russo arxiv

Large-scale video generative models have shown emerging capabilities as zero-shot visual planners, yet video-generated plans often violate temporal consistency and physical constraints, leading to failures when mapped to…

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

2026-02-20 · Zhan Liu, Changli Tang, Yuxin Wang, Zhiyuan Zhu 외 arxiv

Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundamental dimensionality mismatch that preclu…

Spatial ReasoningVisual Grounding

Explainable Semantic Space by Grounding Language to Vision with Cross-Modal Contrastive Learning

2021-11-13 · NeurIPS 2021 12 · Yizhen Zhang, Minkyu Choi, Kuan Han, Zhongming Liu

In natural language processing, most models try to learn semantic representations merely from texts. The learned representations encode the distributional semantics but fail to connect to any knowledge about the physical…

Contrastive LearningImage RetrievalLanguage ModelingLanguage Modelling+1