paper-with-me

홈 › Papers

VISTA: Vision-Grounded and Physics-Validated Adaptation of UMI data for VLA Training

2026-06-03 · Siyuan Yang, Linzheng Guo, Ouyang Lu, Zhaxizhuoma, Daoran Zhang, Xinmiao Wang, Ting Xiao, Fangzheng Yan, Zhijun Chen, Yan Ding, Chao Yu, Chenjia Bai, Xuelong Li arxiv

Universal Manipulation Interface (UMI) enables scalable real-world robot data collection without hardware-specific teleoperation, yet leveraging UMI data to train large-scale Vision-Language-Action (VLA) models remains fundamentally challenging. We identify two critical mismatches: wrist-mounted fisheye views, with severe radial distortion and local gripper-centric perspectives, are out-of-distribution for pretrained VLMs; and human-collected trajectories frequently violate kinematic limits, incur collisions, or exceed controller bandwidth, teaching VLA policies physically infeasible actions. To address the challenges, we present VISTA, a framework that bridges this dual gap through three synergistic components. (i)~UMI-VQA, the first large-scale VQA dataset tailored to wrist-mounted fisheye observations, aligns VLM representations to the distorted visual regime via auxiliary vision-language supervision. (ii)~A systematic physical-validation pipeline performs a data-completeness pre-check and scores each valid trajectory for trajectory continuity, self-collision risk, and execution fidelity before it enters training. (iii)~A two-stage co-training recipe jointly learns vision-language grounding on UMI-VQA and action prediction on validated trajectories. Our experiments empirically show that incorporating UMI-VQA consistently improves downstream policy performance, and that physical-validation scores are strongly predictive of deployment success. On diverse simulation and real-world manipulation tasks, VISTA significantly outperforms strong baselines including $π_{0.5}$, LingBot-VLA, and Wall-X. We release the physical-validation pipeline, UMI-VQA, validated trajectory data, and the pre-trained model for the community.

📄 PDF Abstract BibTeX arXiv:2606.04708

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation

2025-08-08 · Kaiyuan Jiang, Ruoxi Sun, Ying Cao, Yuqi Xu 외 arxiv

We present VISTAR, a user-centric, multi-dimensional benchmark for text-to-image (T2I) evaluation that addresses the limitations of existing metrics. VISTAR introduces a two-tier hybrid paradigm: it employs deterministic…

VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation

2026-08-28 · Zewen Ding, Zezhong Wu, Zhou Tao, Shida Wang 외 arxiv

On-policy self-distillation (OPSD) improves reasoning by training a problem-only student on its own rollouts using dense token-level supervision from a privileged teacher that also sees a reference solution. However, sta…

The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering

2025-02-05 · Zhuowei Li, Haizhou Shi, Yunhe Gao, Di Liu 외

Large Vision-Language Models (LVLMs) can reason effectively over both textual and visual inputs, but they tend to hallucinate syntactically coherent yet visually ungrounded contents. In this paper, we investigate the int…

Hallucination

Unbiased Visual Reasoning with Controlled Visual Inputs

2025-12-19 · Zhaonan Li, Shijie Lu, Fei Wang, Jacob Dineen 외 arxiv

End-to-end Vision-language Models (VLMs) often answer visual questions by exploiting spurious correlations instead of causal visual evidence, and can become more shortcut-prone when fine-tuned. We introduce VISTA (Visual…

Reinforcement LearningVisual Reasoning

VISTA: Visually Inferred Spatial ConTact Attention for Contact-Rich Manipulation

2026-08-26 · Jiayi Chen, Wenlong Dong, Yan Huang, Xianglin Chen 외 arxiv

Contact-rich manipulation requires precise interaction feedback. While vision-centric imitation learning is prevalent, external visual observations provide indirect and ambiguous cues about contact states, particularly u…