paper-with-me

Papers

Trajectory-Constrained Deep Latent Visual Attention for Improved Local Planning in Presence of Heterogeneous Terrain

2021-12-09 · Stefan Wapnick, Travis Manderson, David Meger, Gregory Dudek

We present a reward-predictive, model-based deep learning method featuring trajectory-constrained visual attention for local planning in visual navigation tasks. Our method learns to place visual attention at locations in latent image space which follow trajectories caused by vehicle control actions to enhance predictive accuracy during planning. The attention model is jointly optimized by the task-specific loss and an additional trajectory-constraint loss, allowing adaptability yet encouraging a regularized structure for improved generalization and reliability. Importantly, visual attention is applied in latent feature map space instead of raw image space to promote efficient planning. We validated our model in visual navigation tasks of planning low turbulence, collision-free trajectories in off-road settings and hill climbing with locking differentials in the presence of slippery terrain. Experiments involved randomized procedural generated simulation and real-world environments. We found our method improved generalization and learning efficiency when compared to no-attention and self-attention alternatives.

📄 PDF Abstract BibTeX arXiv:2112.04684

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Navigation

Similar Papers 제목 키워드 기반

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy

2026-05-08 · Zuojin Tang, Shengchao Yuan, Xiaoxin Bai, Zhiyuan Jing 외 arxiv

Vision-language-action (VLA) models increasingly rely on auxiliary world modules to plan over long horizons, yet how such modules should be parameterized on top of a pretrained VLA remains an open design question. Existi…

CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization

2026-05-09 · Ziyang Ding, Linjian Meng, Yiming Wu, Yuhan Li 외 arxiv

Due to the potential for exploratory reasoning of Latent Visual Reasoning, recent works tend to enable MLLMs (Multimodal Large Language Models) to perform visual reasoning by propagating continuous hidden states instead …

Reinforcement LearningVisual Reasoning

A Robust Pavement Mapping System Based on Normal-Constrained Stereo Visual Odometry

2019-10-29 · Huaiyang Huang, Rui Fan, Yilong Zhu, Ming Liu 외

Pavement condition is crucial for civil infrastructure maintenance. This task usually requires efficient road damage localization, which can be accomplished by the visual odometry system embedded in unmanned aerial vehic…

Visual Odometry

LaWM: Least Action World Models for Long-Horizon Physical Consistency from Visual Observations

2026-05-08 · Qixin Xiao, Maani Ghaffari arxiv

Learning predictive world models from visual observations is a core problem in embodied AI, with applications to model-based reinforcement learning and robotic planning. Existing latent world models typically generate fu…

Reinforcement LearningVideo Generation

From Competition to Coopetition: Coopetitive Training-Free Image Editing Based on Text Guidance

2026-04-17 · Jinhao Shen, Haoqian Du, Xulu Zhang, Xiao-Yong Wei 외 arxiv

Text-guided image editing, a pivotal task in modern multimedia content creation, has seen remarkable progress with training-free methods that eliminate the need for additional optimization. Despite recent progress, exist…

Image Editing