paper-with-me

Robot Manipulation

5개 벤치마크 · 논문 826편 · 이 태스크의 논문 보기 →

Benchmarks

CALVIN

결과 23개

RLBench

결과 20개

MimicGen

결과 7개

SimplerEnv-Widow X

결과 7개

Most implemented

Papers

DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot Manipulation

2026-09-09 · Nisarga Nilavadi, Ralf Römer, Moritz Reuss, Michael Krawez 외 arxiv

Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are…

Robot Manipulation

GTA-2: A Multi-VLM Framework for Synthesizing Robot Manipulation Skills via Grounded Task Axes

2026-09-09 · M. Yunus Seker, Shobhit Aggarwal, Ruwan Wickramarachchi, Jonathan Francis 외 arxiv

Robotic manipulation tasks are often decomposed into behaviors or skills. However, one often needs to predefine these behaviors for specific tasks or try to cover a wide range of tasks using generic skills. As a result, …

Robot Manipulation

CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements

2026-09-07 · Hongxiang Zhao, Mutian Xu, Zeyu Jin, Yiming Hao 외 hf

Transferring human hand demonstrations to robotic grippers has recently emerged as a cost-effective solution for robot learning. However, existing methods are largely confined to simple, planar tasks and fail to handle c…

Robot Manipulation

GRAFT: Grounded and Efficient Online Reinforcement Adaptation for Fine-Grained Robot Manipulation

2026-08-27 · Yibo Qiu, Haoliang Ye, Shu'ang Sun, Zan Huang 외 arxiv

Pretrained vision-language-action (VLA) policies provide strong priors for robot manipulation, yet adapting them online to fine-grained biomedical tasks remains challenging. Task success often hinges on subtle, view-depe…

Robot ManipulationVisual Grounding

Riemann-1.0: An Embodied World Action Model for Physical AI

2026-08-27 · Haofeng Sun, Jiangbo Pei, Fei Kang, Zexiang Liu 외 arxiv

We introduce Riemann-1.0, a fully causal autoregressive World Action Model for embodied intelligence. Riemann-1.0 jointly models multi-view visual observations, robot states, and embodiment-specific actions within a unif…

Robot Manipulation

TemporalFlow-VLA: Learning Physically Grounded Execution History for Long-Horizon Robot Manipulation

2026-08-27 · Jiarui Yang, Yehao Lu, Yuning Su, Yu Zhong 외 arxiv

Vision-language-action (VLA) models leverage pretrained vision-language representations for robot control, yet simply adding historical frames does not reliably capture recent physical change. This is especially problema…

Robot Manipulation

전체 826편 보기 →