paper-with-me

홈 › Papers

NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards

2025-11-18 · Chia-Yu Hung, Navonil Majumder, Haoyuan Deng, Liu Renhang, Yankang Ang, Amir Zadeh, Chuan Li, Dorien Herremans, Ziwei Wang, Soujanya Poria arxiv

Vision--language--action (VLA) models have recently shown promising performance on a variety of embodied tasks, yet they still fall short in reliability and generalization, especially when deployed across different embodiments or real-world environments. In this work, we introduce NORA-1.5, a VLA model built from the pre-trained NORA backbone by adding to it a flow-matching-based action expert. This architectural enhancement alone yields substantial performance gains, enabling NORA-1.5 to outperform NORA and several state-of-the-art VLA models across both simulated and real-world benchmarks. To further improve robustness and task success, we develop a set of reward models for post-training VLA policies. Our rewards combine (i) an action-conditioned world model (WM) that evaluates whether generated actions lead toward the desired goal, and (ii) a deviation-from-ground-truth heuristic that distinguishes good actions from poor ones. Using these reward signals, we construct preference datasets and adapt NORA-1.5 to target embodiments through direct preference optimization (DPO). Extensive evaluations show that reward-driven post-training consistently improves performance in both simulation and real-robot settings, demonstrating significant VLA model-reliability gains through simple yet effective reward models. Our findings highlight NORA-1.5 and reward-guided post-training as a viable path toward more dependable embodied agents suitable for real-world deployment.

📄 PDF Abstract BibTeX arXiv:2511.14659

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming

2025-08-04 · Shuo Wang, Yongcai Wang, Zhaoxin Fan, Yucheng Wang 외 arxiv

Vision-Language Navigation (VLN) tasks often leverage panoramic RGB and depth inputs to provide rich spatial cues for action planning, but these sensors can be costly or less accessible in real-world deployments. Recent …

Vision-Language Navigation

Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation

2024-06-14 · Zihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu 외

Vision-and-language navigation (VLN) enables the agent to navigate to a remote location in 3D environments following the natural language instruction. In this field, the agent is usually trained and evaluated in the navi…

NavigateVision and Language Navigation

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

2025-04-28 · Chia-Yu Hung, Qi Sun, Pengfei Hong, Amir Zadeh 외

Existing Visual-Language-Action (VLA) models have shown promising performance in zero-shot scenarios, demonstrating impressive task execution and reasoning capabilities. However, a significant challenge arises from the l…

Task PlanningVision-Language-ActionVisual Reasoning

Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation

2025-11-02 · Xiangyu Shi, Zerui Li, Yanyuan Qiao, Qi Wu arxiv

Recent advances in Vision-and-Language Navigation in Continuous Environments (VLN-CE) have leveraged multimodal large language models (MLLMs) to achieve zero-shot navigation. However, existing methods often rely on panor…

SAP: Segment Any 4K Panorama

2026-03-13 · Lutao Jiang, Zidong Cao, Weikai Chen, Xu Zheng 외 arxiv

Promptable instance segmentation is widely adopted in embodied and AR systems, yet the performance of foundation models trained on perspective imagery often degrades on 360° panoramas. In this paper, we introduce Segment…

Instance SegmentationVideo Segmentation