paper-with-me

Papers

Joint Perception and Control as Inference with an Object-based Implementation

2019-03-04 · Minne Li, Zheng Tian, Pranav Nashikkar, Ian Davies, Ying Wen, Jun Wang

Existing model-based reinforcement learning methods often study perception modeling and decision making separately. We introduce joint Perception and Control as Inference (PCI), a general framework to combine perception and control for partially observable environments through Bayesian inference. Based on the fact that object-level inductive biases are critical in human perceptual learning and reasoning, we propose Object-based Perception Control (OPC), an instantiation of PCI which manages to facilitate control using automatic discovered object-based representations. We develop an unsupervised end-to-end solution and analyze the convergence of the perception model update. Experiments in a high-dimensional pixel environment demonstrate the learning effectiveness of our object-based perception control approach. Specifically, we show that OPC achieves good perceptual grouping quality and outperforms several strong baselines in accumulated rewards.

📄 PDF Abstract BibTeX arXiv:1903.01385

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian InferenceDecision MakingModel-based Reinforcement LearningObject

Similar Papers 제목 키워드 기반

The Components of Collaborative Joint Perception and Prediction -- A Conceptual Framework

2025-01-27 · Lei Wan, Hannan Ejaz Keen, Alexey Vinel

Connected Autonomous Vehicles (CAVs) benefit from Vehicle-to-Everything (V2X) communication, which enables the exchange of sensor data to achieve Collaborative Perception (CP). To reduce cumulative errors in perception m…

Autonomous Vehiclesmotion predictionPrediction

The free energy principle for action and perception: A mathematical review

2017-05-24

The 'free energy principle' (FEP) has been suggested to provide a unified theory of the brain, integrating data and theory relating to action, perception, and learning. The theory and implementation of the FEP combines i…

Learning Theory

HENet: Hybrid Encoding for End-to-end Multi-task 3D Perception from Multi-view Cameras

2024-04-03 · Zhongyu Xia, Zhiwei Lin, Xinhao Wang, Yongtao Wang 외

Three-dimensional perception from multi-view cameras is a crucial component in autonomous driving systems, which involves multiple tasks like 3D object detection and bird's-eye-view (BEV) semantic segmentation. To improv…

3D Object DetectionAutonomous Drivingobject-detectionObject Detection+1

dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought

2025-09-30 · Junjie Wen, Minjie Zhu, Jiaming Liu, Zhiyuan Liu 외 arxiv

Vision-Language-Action (VLA) models are emerging as a next-generation paradigm for robotics. We introduce dVLA, a diffusion-based VLA that leverages a multimodal chain-of-thought to unify visual perception, language reas…

Behavior Cloning for Active Perception with Low-Resolution Egocentric Vision

2026-05-13 · Anthony Bilic, Chen Chen, Ladislau Bölöni arxiv

We investigate whether behavior cloning is sufficient to produce active perception in a structured object-finding task. A low-cost robot arm equipped with a wrist-mounted egocentric RGB camera must reposition to center a…