paper-with-me

홈 › Papers

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

2026-06-12 · He Zhang, Lingzhu Xiang, Haitao Lin, Zeyu Huang, Minghui Wang, Dingyan Zhong, Yubo Dong, Yihao Wu, Yongming Rao, Dongsheng Zhang, Wanjia He, Ling Chen, Kai Huang, Jiahao Chen, Sichang Su, Xumin Yu, Ziyi Wang, Chengwei Zhu, Xiao Teng, Yuchun Guo, Yufeng Zhang, Yuandong Liu, Rui Wang, Zisheng Lu, Han Hu, Zhengyou Zhang arxiv

In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: data collection, model design, continued pre-training and supervised fine-tuning, RL post-training, and real-world deployment. Each component serves a distinct role in this stack.

📄 PDF Abstract BibTeX arXiv:2606.14409

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

3D-VLA: A 3D Vision-Language-Action Generative World Model

2024-03-14 · Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang 외

Recent vision-language-action (VLA) models rely on 2D inputs, lacking integration with the broader realm of the 3D physical world. Furthermore, they perform action prediction by learning a direct mapping from perception …

Language ModellingLarge Language Modelmultimodal generationVision-Language-Action

PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models

2026-03-31 · Amirreza Rouhi, Parikshit Sakurikar, Satya Sai Reddy, Narsimha Menga 외 arxiv

A critical gap exists between the general-purpose visual understanding of state-of-the-art physical AI models and the specialized perceptual demands of structured real-world deployment environments. We present PRISM, a 2…

Action Understanding

Perceptual Quality Assessment for Embodied AI

2025-05-22 · Chunyi Li, Jiaohao Xiao, Jianbo Zhang, Farong Wen 외

Embodied AI has developed rapidly in recent years, but it is still mainly deployed in laboratories, with various distortions in the Real-world limiting its application. Traditionally, Image Quality Assessment (IQA) metho…

Image Quality AssessmentVision-Language-Action

Survey of Vision-Language-Action Models for Embodied Manipulation

2025-08-21 · Haoran Li, Yuhui Chen, Wenbo Cui, Weiheng Liu 외 arxiv

Embodied intelligence systems, which enhance agent capabilities through continuous environment interactions, have garnered significant attention from both academia and industry. Vision-Language-Action models, inspired by…

Embodied Image Compression

2025-12-12 · Chunyi Li, Rui Qing, Jianbo Zhang, Yuan Tian 외 arxiv

Image Compression for Machines (ICM) has emerged as a pivotal research direction in the field of visual data compression. However, with the rapid evolution of machine intelligence, the target of compression has shifted f…

Image Compression