paper-with-me

Papers

Learning Visual Predictive Models of Physics for Playing Billiards

2015-11-23 · Katerina Fragkiadaki, Pulkit Agrawal, Sergey Levine, Jitendra Malik

The ability to plan and execute goal specific actions in varied, unexpected settings is a central requirement of intelligent agents. In this paper, we explore how an agent can be equipped with an internal model of the dynamics of the external world, and how it can use this model to plan novel actions by running multiple internal simulations ("visual imagination"). Our models directly process raw visual input, and use a novel object-centric prediction formulation based on visual glimpses centered on objects (fixations) to enforce translational invariance of the learned physical laws. The agent gathers training data through random interaction with a collection of different environments, and the resulting model can then be used to plan goal-directed actions in novel environments that the agent has not seen before. We demonstrate that our agent can accurately plan actions for playing a simulated billiards game, which requires pushing a ball into a target position or into collision with another ball.

📄 PDF Abstract BibTeX arXiv:1511.07404

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CueTip: An Interactive and Explainable Physics-aware Pool Assistant

2025-01-30 · Sean Memery, Kevin Denamganai, Jiaxin Zhang, Zehai Tu 외

We present an interactive and explainable automated coaching assistant called CueTip for a variant of pool/billiards. CueTip's novelty lies in its combination of three features: a natural-language interface, an ability t…

Scan, Materialize, Simulate: A Generalizable Framework for Physically Grounded Robot Planning

2025-05-20 · Amine Elhafsi, Daniel Morton, Marco Pavone

Autonomous robots must reason about the physical consequences of their actions to operate effectively in unstructured, real-world environments. We present Scan, Materialize, Simulate (SMS), a unified framework that combi…

Semantic Segmentation

A Perceptual Alphabet for the 10-dimensional Phonetic-prosodic Space

2013-06-11 · Elaine Y L Tsiang

We define an alphabet, the IHA, of the 10-D phonetic-prosodic space. The dimensions of this space are perceptual observables, rather than articulatory specifications. Speech is defined as a random chain in time of the 4-…

HiLight: Technical Report on the Motern AI Video Language Model

2024-07-10 · Zhiting Wang, Qiangong Zhou, Kangjie Yang, Zongyang Liu 외

This technical report presents the implementation of a state-of-the-art video encoder for video-text modal alignment and a video conversation framework called HiLight, which features dual visual towers. The work is divid…

Language ModelingLanguage Modelling

Piecewise-constant Neural ODEs

2021-06-11 · Sam Greydanus, Stefan Lee, Alan Fern

Neural networks are a popular tool for modeling sequential data but they generally do not treat time as a continuous variable. Neural ODEs represent an important exception: they parameterize the time derivative of a hidd…