paper-with-me

Papers

Intrinsic Dynamics-Driven Generalizable Scene Representations for Vision-Oriented Decision-Making Applications

2024-05-30 · Dayang Liang, Jinyang Lai, Yunlong Liu

How to improve the ability of scene representation is a key issue in vision-oriented decision-making applications, and current approaches usually learn task-relevant state representations within visual reinforcement learning to address this problem. While prior work typically introduces one-step behavioral similarity metrics with elements (e.g., rewards and actions) to extract task-relevant state information from observations, they often ignore the inherent dynamics relationships among the elements that are essential for learning accurate representations, which further impedes the discrimination of short-term similar task/behavior information in long-term dynamics transitions. To alleviate this problem, we propose an intrinsic dynamics-driven representation learning method with sequence models in visual reinforcement learning, namely DSR. Concretely, DSR optimizes the parameterized encoder by the state-transition dynamics of the underlying system, which prompts the latent encoding information to satisfy the state-transition process and then the state space and the noise space can be distinguished. In the implementation and to further improve the representation ability of DSR on encoding similar tasks, sequential elements' frequency domain and multi-step prediction are adopted for sequentially modeling the inherent dynamics. Finally, experimental results show that DSR has achieved significant performance improvements in the visual Distracting DMControl control tasks, especially with an average of 78.9\% over the backbone baseline. Further results indicate that it also achieves the best performances in real-world autonomous driving applications on the CARLA simulator. Moreover, qualitative analysis results validate that our method possesses the superior ability to learn generalizable scene representations on visual tasks. The source code is available at https://github.com/DMU-XMU/DSR.

📄 PDF Abstract BibTeX arXiv:2405.19736

Code (1)

dmu-xmu/dsr 공식 구현 pytorch

Tasks

Autonomous DrivingDecision MakingRepresentation Learning

Methods 이 논문이 사용한 방법론

Entropy Regularization 설명 없음
PPO Proximal Policy Optimization, or PPO, is a policy gradient method for reinforcement learning. The motivation was to have an algorithm with the data efficiency and reliable…
CARLA CARLA is an open-source simulator for autonomous driving research. CARLA has been developed from the ground up to support development, training, and validation of autonomous urban…

Similar Papers 제목 키워드 기반

DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous Driving

2025-05-25 · Chen Shi, Shaoshuai Shi, Kehua Sheng, Bo Zhang 외

Data-driven learning has advanced autonomous driving, yet task-specific models struggle with out-of-distribution scenarios due to their narrow optimization objectives and reliance on costly annotated data. We present Dri…

Autonomous DrivingImage GenerationRepresentation LearningWorld Knowledge

Measuring and Modeling Physical Intrinsic Motivation

2023-05-22 · Julio Martinez, Felix Binder, Haoliang Wang, Nick Haber 외

Humans are interactive agents driven to seek out situations with interesting physical dynamics. Here we formalize the functional form of physical intrinsic motivation. We first collect ratings of how interesting humans f…

Prediction

G3R: Gradient Guided Generalizable Reconstruction

2024-09-28 · Yun Chen, Jingkang Wang, Ze Yang, Sivabalan Manivasagam 외

Large scale 3D scene reconstruction is important for applications such as virtual reality and simulation. Existing neural rendering approaches (e.g., NeRF, 3DGS) have achieved realistic reconstructions on large scenes, b…

3DGS3D Scene ReconstructionNeRFNeural Rendering

Boosting Generalizable Depth Estimation in Endoscopy by Mixture of Lightweight Experts and Intrinsic Image Alignment

2026-08-01 · Liangjing Shao, Beilei Cui, Yiming Huang, Changjing Liu 외 arxiv

Depth estimation is a significant task for 3D perception in endoscopic surgeries. However, illumination interference and feature diversity in various endoscopic scenes are still challenges for generalizable depth estimat…

parameter-efficient fine-tuningDepth Estimation

NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition

2026-04-17 · Junguang Yao, Wenye Liu, Stjepan Picek, Yue Zheng arxiv

Visual speaker recognition based on lip motion offers a silent, hands-free, and behavior-driven biometric solution that remains effective even when acoustic cues are unavailable. Compared to traditional methods that rely…

Speaker Recognition