paper-with-me

홈 › Papers

PhysV2A: Reachability-Gated and Semantic-Mask-Constrained Feasibility Completion for Video-to-Robot Manipulation

2026-07-10 · Haohui Huang, Junda Duan, Tao Teng, Chenguang Yang arxiv

Video-based manipulation provides object-centric motion priors from human demonstrations, generated videos, or RGB-D observations, but such priors are typically embodiment-agnostic and cannot be directly executed by a specific robot. This paper presents \textbf{PhysV2A}, a reachability-gated and semantic-mask-constrained feasibility-completion framework for converting video-derived 6D object motion into robot-executable manipulation trajectories. The key idea is to treat grasp feasibility as trajectory-conditioned rather than local: each RGB-D-generated 6-DoF grasp candidate is rigidly coupled with the recovered object motion to form a grasp-conditioned TCP trajectory hypothesis. PhysV2A then performs hierarchical reachability-gated selection, where infeasible grasp--trajectory pairs are rejected by robot-centric kinematic checks and surviving candidates are ranked by downstream execution suitability. For the selected reachable trajectory, a VLM-assisted and rule-validated S-Mask identifies task-critical and relaxable Cartesian components, enabling semantic-mask-constrained manipulability refinement through redundancy-first optimization and bounded Cartesian relaxation. Real-robot experiments on four tabletop manipulation tasks show that PhysV2A improves task success over representative video-prior and IK-only baselines, reduces kinematic-feasibility failures, and produces better-conditioned trajectories with bounded semantic deviations.

📄 PDF Abstract BibTeX arXiv:2607.09365

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability

2025-03-11 · CVPR 2025 1 · Weijie Zhou, Manli Tao, Chaoyang Zhao, Haiyun Guo 외

Understanding the environment and a robot's physical reachability is crucial for task execution. While state-of-the-art vision-language models (VLMs) excel in environmental perception, they often generate inaccurate or i…

Visual Reasoning

MMPhysVideo: Physically Plausible Video Generation Through Joint RGB-Perception Modeling

2026-04-03 · Shubo Lin, Xuanyang Zhang, Wei Cheng, Weiming Hu 외 arxiv

Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel-only reconstruction. To address this, we propose MMPhysVideo, the first…

Video Generation

Reachability Analysis Using Constrained Polynomial Logical Zonotopes

2024-03-27 · Ahmad Hafez, Frank J. Jiang, Karl H. Johansson, Amr Alanwar

In this paper, we propose reachability analysis using constrained polynomial logical zonotopes. We perform reachability analysis to compute the set of states that could be reached. To do this, we utilize a recently intro…

Computational Efficiency

Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation

2026-05-13 · Qisong He, Xinmiao Huang, Jinwei Hu, Zhuoyun Li 외 arxiv

Safe navigation for mobile robots demands policies that remain reliable under the high-consequence perception uncertainty of cluttered environments. Yet most existing safe reinforcement learning (RL) methods assess safet…

Reinforcement LearningRobot Navigation

PhysVid: Physics Aware Local Conditioning for Generative Video Models

2026-03-27 · Saurabh Pathak, Elahe Arani, Mykola Pechenizkiy, Bahram Zonooz arxiv

Generative video models achieve high visual fidelity but often violate basic physical principles, limiting reliability in real-world settings. Prior attempts to inject physics rely on conditioning: frame-level signals ar…