paper-with-me

Papers Robot Manipulation

“Robot Manipulation” 태그가 달린 논문 826편 · 필터 해제

DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation

2026-09-09 · Nisarga Nilavadi, Ralf Römer, Moritz Reuss, Michael Krawez 외 arxiv

Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are…

Robot Manipulation

GTA-2: A Multi-VLM Framework for Synthesizing Robot Manipulation Skills via Grounded Task Axes

2026-09-09 · M. Yunus Seker, Shobhit Aggarwal, Ruwan Wickramarachchi, Jonathan Francis 외 arxiv

Robotic manipulation tasks are often decomposed into behaviors or skills. However, one often needs to predefine these behaviors for specific tasks or try to cover a wide range of tasks using generic skills. As a result, …

Robot Manipulation

CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements

2026-09-07 · Hongxiang Zhao, Mutian Xu, Zeyu Jin, Yiming Hao 외 hf

Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle c…

Robot Manipulation

GRAFT: Grounded and Efficient Online Reinforcement Adaptation for Fine-Grained Robot Manipulation

2026-08-27 · Yibo Qiu, Haoliang Ye, Shu'ang Sun, Zan Huang 외 arxiv

Pretrained vision-language-action (VLA) policies provide strong priors for robot manipulation, yet adapting them online to fine-grained biomedical tasks remains challenging. Task success often hinges on subtle, view-depe…

Robot ManipulationVisual Grounding

Riemann-1.0: An Embodied World Action Model for Physical AI

2026-08-27 · Haofeng Sun, Jiangbo Pei, Fei Kang, Zexiang Liu 외 arxiv

We introduce Riemann-1.0, a fully causal autoregressive World Action Model for embodied intelligence. Riemann-1.0 jointly models multi-view visual observations, robot states, and embodiment-specific actions within a unif…

Robot Manipulation

TemporalFlow-VLA: Learning Physically Grounded Execution History for Long-Horizon Robot Manipulation

2026-08-27 · Jiarui Yang, Yehao Lu, Yuning Su, Yu Zhong 외 arxiv

Vision-language-action (VLA) models leverage pretrained vision-language representations for robot control, yet simply adding historical frames does not reliably capture recent physical change. This is especially problema…

Robot Manipulation

PredVLA: A Sub-Million-Parameter Predictive-Coding Policy for Robot Manipulation

2026-08-27 · Hiroki Sawada, Shunichi Kasahara arxiv

Large pretrained vision-language-action models dominate modern robot-manipulation benchmarks, but it remains unclear how much model scale is necessary for strong language-conditioned control, or whether fundamentally dif…

Robot Manipulation

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

2026-08-26 · Zhe Liu, Jinghua Hou, Yuxiang Lu, Zhenya Yang 외 arxiv

Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting their ability to retain past observatio…

Robot Manipulation

LM-X: Explainable Vision--Language--Action Modeling via Progress, Event, and Uncertainty Prediction

2026-08-26 · Jin Lou, Jingxuan Zhu, Andong Chen, Xupeng Wang 외 arxiv

Large-scale vision--language--action (VLA) policies have advanced generalist robot control, yet most remain stimulus-to-action black boxes: actions are exposed, but their explanatory state is not. They provide no native …

Robot Manipulation

InstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation

2026-08-24 · Mengao Zhao, Ziang Li, Chaodong Huang, Mengchen Ma 외 arxiv

Vision-language-action (VLA) models have made general-purpose robot manipulation increasingly plausible by conditioning robot actions on natural-language instructions. A key test of such generality is whether policies ac…

Instruction FollowingRobot ManipulationSpatial Reasoning

Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

2026-08-21 · Varun Giridhar, Anant Khandelwal, Jeremy A. Collins, Ignat Georgiev 외 arxiv

Behaviour Cloning (BC) has driven remarkable progress in robot manipulation, yet it is fundamentally limited by its inability to self-improve: a policy that fails cannot learn from that failure without additional human d…

Reinforcement LearningRobot Manipulation

DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

2026-08-20 · Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao 외 arxiv

Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-sight gaps. Existing single-frame and …

Robot Manipulation

What Matters for Latent Actions in Robot Learning

2026-08-20 · Xizhou Bu, Qingda Hu, Lei Zhou, Lingfeng Zhang 외 arxiv

Latent Action Models (LAMs) have emerged as a promising paradigm for enabling robot learning to leverage large-scale unlabeled videos through latent actions that serve as compact surrogates for physical actions. Despite …

Robot Manipulation

HiTac-WAM: A Hierarchical Tactile World Action Model for Contact-Rich Robot Manipulation

2026-08-20 · Chao Xue, Chaofan Zhang, Wenxuan Ma, Guocai Yao 외 arxiv

World action models jointly predict future visual observations and actions, whereas existing tactile-aware variants typically represent future touch as an image or latent stream without modeling the physical dependencies…

Robot Manipulation

Dream2Reward: Transition-Alignment Reward Models from Positive Demonstrations for Robotic Manipulation

2026-08-19 · Haoyu Zhang, Zecui Zeng, Bin Wang, Lusong Li 외 arxiv

Learning robotic policies requires dense rewards that remain informative when behavior departs from successful demonstrations. Progress-based rewards estimate how far an observation has advanced along a nominal successfu…

Robot Manipulation

CompCPZ: Preserving Multi-Modal Intent in Language-Guided Robot Manipulation

2026-08-18 · Zhen Zhang, Ahmad Hafez, Peng Xie, Yanliang Huang 외 arxiv

A robot asked to "place the cup near the red plate or the blue plate" may reach the centroid between them and appear geometrically successful, while satisfying neither disjunct of the instruction. This silent semantic fa…

Robot Manipulation

ORPA: Online Residual Policy Adaptation for Robot Manipulation Control with Human Feedback

2026-08-18 · Muhammad A. Muttaqien, Tomohiro Motoda, Ryo Hanai, Yukiyasu Domae arxiv

Robotic manipulation policies trained via imitation learning, such as Action Chunking with Transformers (ACT), can achieve strong performance under ideal conditions but often remain sensitive to small execution errors an…

Robot Manipulation

Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory

2026-08-17 · Bingxin Xu, Yuzhang Shang, Emilio Ferrara arxiv

Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) models increasingly master the individual skills, yet the chain still fails: errors compound beyond …

Robot Manipulation

τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

2026-08-17 · Xiaowei Cai, Yunuo Cai, Bingao Chen, Jingxiao Chen 외 arxiv

Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision …

Robot Manipulation

Unified Condition-Action Modeling for Accurate One-Step Action Generation

2026-08-17 · Xinyu Zhou, Zikun Cai, Kuangji Zuo, Gen Li 외 arxiv

Robot manipulation requires policies that are both accurate and efficient, as robot control must respond to changing observations under tight latency constraints. Recent diffusion and flow policies are promising, but the…

Representation LearningRobot Manipulation
1–20 / 826 다음 →