paper-with-me

Papers

ShapeStacks: Learning Vision-Based Physical Intuition for Generalised Object Stacking

2018-04-21 · ECCV 2018 9 · Oliver Groth, Fabian B. Fuchs, Ingmar Posner, Andrea Vedaldi

Physical intuition is pivotal for intelligent agents to perform complex tasks. In this paper we investigate the passive acquisition of an intuitive understanding of physical principles as well as the active utilisation of this intuition in the context of generalised object stacking. To this end, we provide: a simulation-based dataset featuring 20,000 stack configurations composed of a variety of elementary geometric primitives richly annotated regarding semantics and structural stability. We train visual classifiers for binary stability prediction on the ShapeStacks data and scrutinise their learned physical intuition. Due to the richness of the training data our approach also generalises favourably to real-world scenarios achieving state-of-the-art stability prediction on a publicly available benchmark of block towers. We then leverage the physical intuition learned by our model to actively construct stable stacks and observe the emergence of an intuitive notion of stackability - an inherent object affordance - induced by the active stacking task. Our approach performs well even in challenging conditions where it considerably exceeds the stack height observed during training or in cases where initially unstable structures must be stabilised via counterbalancing.

📄 PDF Abstract BibTeX arXiv:1804.08018

Code (1)

ogroth/shapestacks 공식 구현 tf

Tasks

Physical Intuition

Similar Papers 제목 키워드 기반

RELATE: Physically Plausible Multi-Object Scene Synthesis Using Structured Latent Spaces

2020-07-02 · NeurIPS 2020 12 · Sebastien Ehrhardt, Oliver Groth, Aron Monszpart, Martin Engelcke 외

We present RELATE, a model that learns to generate physically plausible scenes and videos of multiple interacting objects. Similar to other generative approaches, RELATE is trained end-to-end on raw, unlabeled data. RELA…

ObjectScene Generation

Can Vision Language Models Learn Intuitive Physics from Interaction?

2026-02-05 · Luca M. Schulze Buschoff, Konstantinos Voudouris, Can Demircan, Eric Schulz arxiv

Pre-trained vision language models do not have good intuitions about the physical world. Recent work has shown that supervised fine-tuning can improve model performance on simple physical tasks. However, fine-tuned model…

Reinforcement Learning

WoW: Towards a World omniscient World model Through Embodied Interaction

2025-09-26 · Xiaowei Chi, Peidong Jia, Chun-Kai Fan, Xiaozhu Ju 외 arxiv

Humans develop an understanding of intuitive physics through active interaction with the world. This approach is in stark contrast to current video models, such as Sora, which rely on passive observation and therefore st…

Physical Intuition

ICPRL: Acquiring Physical Intuition from Interactive Control

2026-03-01 · Xinrun Xu, Pi Bu, Ye Wang, Börje F. Karlsson 외 arxiv

VLMs excel at static perception but falter in interactive reasoning in dynamic physical environments, which demands planning and adaptation to dynamic outcomes. Existing physical reasoning methods often depend on abstrac…

Reinforcement LearningPhysical Intuition

A Link Quality Model for Generalised Frequency Division Multiplexing

2017-10-25

5G systems aim to achieve extremely high data rates, low end-to-end latency and ultra-low power consumption. Recently, there has been considerable interest in the design of 5G physical layer waveforms. One important cand…