paper-with-me

홈 › Papers

VT-Refine: Learning Bimanual Assembly with Visuo-Tactile Feedback via Simulation Fine-Tuning

2025-10-16 · Binghao Huang, Jie Xu, Iretiayo Akinola, Wei Yang, Balakumar Sundaralingam, Rowland O'Flaherty, Dieter Fox, Xiaolong Wang, Arsalan Mousavian, Yu-Wei Chao, Yunzhu Li arxiv

Humans excel at bimanual assembly tasks by adapting to rich tactile feedback -- a capability that remains difficult to replicate in robots through behavioral cloning alone, due to the suboptimality and limited diversity of human demonstrations. In this work, we present VT-Refine, a visuo-tactile policy learning framework that combines real-world demonstrations, high-fidelity tactile simulation, and reinforcement learning to tackle precise, contact-rich bimanual assembly. We begin by training a diffusion policy on a small set of demonstrations using synchronized visual and tactile inputs. This policy is then transferred to a simulated digital twin equipped with simulated tactile sensors and further refined via large-scale reinforcement learning to enhance robustness and generalization. To enable accurate sim-to-real transfer, we leverage high-resolution piezoresistive tactile sensors that provide normal force signals and can be realistically modeled in parallel using GPU-accelerated simulation. Experimental results show that VT-Refine improves assembly performance in both simulation and the real world by increasing data diversity and enabling more effective policy fine-tuning. Our project page is available at https://binghao-huang.github.io/vt_refine/.

📄 PDF Abstract BibTeX arXiv:2510.14930

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Learning Visuotactile Skills with Two Multifingered Hands

2024-04-25 · Toru Lin, Yu Zhang, Qiyang Li, Haozhi Qi 외

Aiming to replicate human-like dexterity, perceptual experiences, and motion patterns, we explore learning from human demonstrations using a bimanual system with multifingered hands and visuotactile data. Two significant…

TAMEn: Tactile-Aware Manipulation Engine for Closed-Loop Data Collection in Contact-Rich Tasks

2026-04-08 · Longyan Wu, Jieji Ren, Chenghang Jiang, Junxi Zhou 외 arxiv

Handheld paradigms offer an efficient and intuitive way for collecting large-scale demonstration of robot manipulation. However, achieving contact-rich bimanual manipulation through these methods remains a pivotal challe…

Robot Manipulation

TacCoRL: Integrating Tactile Feedback into VLA via Simulation

2026-06-10 · Siyu Ma, Yuqi Liang, Chang Yu, Yunuo Chen 외 arxiv

Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state required for contact-rich tasks. We present …

Reinforcement LearningRobot Manipulation

OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation

2026-03-19 · Yuhang Zheng, Songen Gu, Weize Li, Yupeng Zheng 외 arxiv

Contact-rich manipulation tasks, such as wiping and assembly, require accurate perception of contact forces, friction changes, and state transitions that cannot be reliably inferred from vision alone. Despite growing int…

TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance

2026-01-28 · Zhemeng Zhang, Jiahua Ma, Xincheng Yang, Xin Wen 외 arxiv

Fine-grained and contact-rich manipulation remain challenging for robots, largely due to the underutilization of tactile feedback. To address this, we introduce TouchGuide, a novel cross-policy visuo-tactile fusion parad…

Contrastive Learning