BiFold: Bimanual Cloth Folding with Language Guidance
Cloth folding is a complex task due to the inevitable self-occlusions of clothes, their complicated dynamics, and the disparate materials, geometries, and textures that garments can have. In this work, we learn folding actions conditioned on text commands. Translating high-level, abstract instructions into precise robotic actions requires sophisticated language understanding and manipulation capabilities. To do that, we leverage a pre-trained vision-language model and repurpose it to predict manipulation actions. Our model, BiFold, can take context into account and achieves state-of-the-art performance on an existing language-conditioned folding benchmark. To address the lack of annotated bimanual folding data, we introduce a novel dataset with automatically parsed actions and language-aligned instructions, enabling better learning of text-conditioned manipulation. BiFold attains the best performance on our dataset and demonstrates strong generalization to new instructions, garments, and environments.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingTAGSimilar Papers 제목 키워드 기반
Beyond Static Perception: Integrating Temporal Context into VLMs for Cloth Folding
Manipulating clothes is challenging due to their complex dynamics, high deformability, and frequent self-occlusions. Garments exhibit a nearly infinite number of configurations, making explicit state representations diff…
State EstimationLearning Bimanual Cloth Manipulation with Vision-based Tactile Sensing via Single Robotic Arm
Robotic cloth manipulation remains challenging due to the high-dimensional state space of fabrics, their deformable nature, and frequent occlusions that limit vision-based sensing. Although dual-arm systems can mitigate …
Pose EstimationMonoDuo: Using One Robot Arm to Learn Bimanual Policies
Bimanual coordination is essential for many real-world manipulation tasks, yet learning bimanual robot policies is limited by the scarcity of bimanual robots and datasets. Single-arm robots, however, are widely available…
Point Cloud SegmentationHand Pose EstimationFabricFlowNet: Bimanual Cloth Manipulation with a Flow-based Policy
We address the problem of goal-directed cloth manipulation, a challenging task due to the deformability of cloth. Our insight is that optical flow, a technique normally used for motion estimation in video, can also provi…
Motion EstimationOptical Flow EstimationTaming VR Teleoperation and Learning from Demonstration for Multi-Task Bimanual Table Service Manipulation
This technical report presents the champion solution of the Table Service Track in the ICRA 2025 What Bimanuals Can Do (WBCD) competition. We tackled a series of demanding tasks under strict requirements for speed, preci…