paper-with-me

홈 › Papers

PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention

2025-12-03 · Ziwen Li, Xin Wang, Hanlue Zhang, Runnan Chen, Runqi Lin, Xiao He, Han Huang, Yandong Guo, Fakhri Karray, Tongliang Liu, Mingming Gong arxiv

The Vision-Language-Action (VLA) models have demonstrated remarkable performance on embodied tasks and shown promising potential for real-world applications. However, current VLAs still struggle to produce consistent and precise target-oriented actions, as they often generate redundant or unstable motions along trajectories, limiting their applicability in time-sensitive scenarios.In this work, we attribute these redundant actions to the spatially uniform perception field of existing VLAs, which causes them to be distracted by target-irrelevant objects, especially in complex environments.To address this issue, we propose an efficient PosA-VLA framework that anchors visual attention via pose-conditioned supervision, consistently guiding the model's perception toward task-relevant regions. The pose-conditioned anchor attention mechanism enables the model to better align instruction semantics with actionable visual cues, thereby improving action generation precision and efficiency. Moreover, our framework adopts a lightweight architecture and requires no auxiliary perception modules (e.g., segmentation or grounding networks), ensuring efficient inference. Extensive experiments verify that our method executes embodied tasks with precise and time-efficient behavior across diverse robotic manipulation benchmarks and shows robust generalization in a variety of challenging environments.

📄 PDF Abstract BibTeX arXiv:2512.03724

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

$ω$-EVA: Envision, Verify, and Act with Latent Interactive World Models

2026-06-08 · Zhenguo Sun, Yu Sun, Hande Huang, Alois Knoll arxiv

Embodied policies typically map current observations directly to actions, leaving candidate-action consequences implicit. World models provide predictive supervision, representations, or external simulation, but rarely l…

Proposal-Conditioned Latent Diffusion for Closed-Loop Traffic Scenario Generation

2026-06-25 · Shubham Vaijanath Phoolari, Aleyna Kara, Christoph Lauer, Steven Peters arxiv

Closed-loop traffic simulation remains challenging because it must generate interactive multi-agent behaviors that are scene-consistent and controllable throughout rollout. Prior diffusion-based approaches achieve strong…

HYPE: Hybrid Planning with Ego Proposal-Conditioned Predictions

2025-10-14 · Hang Yu, Julian Jordan, Julian Schmidt, Silvan Lindner 외 arxiv

Safe and interpretable motion planning in complex urban environments needs to reason about bidirectional multi-agent interactions. This reasoning requires to estimate the costs of potential ego driving maneuvers. Many ex…

Motion Planning

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning

2026-06-23 · Xirui Li, Zhe Liu, Xiaoqing Ye, Wenhua Han 외 arxiv

Multimodal driving planning faces a long-standing tension between two paradigms: scoring-based methods benefit from dense reward supervision but are confined to a fixed action vocabulary, while anchor-based methods gener…

BSN: Boundary Sensitive Network for Temporal Action Proposal Generation

2018-06-08 · ECCV 2018 9 · Tianwei Lin, Xu Zhao, Haisheng Su, Chongjing Wang 외

Temporal action proposal generation is an important yet challenging problem, since temporal proposals with rich action content are indispensable for analysing real-world videos with long duration and high proportion irre…

Action DetectionTemporal Action LocalizationTemporal Action Proposal Generation