paper-with-me

Papers

RoTri-Diff: A Spatial Robot-Object Triadic Interaction-Guided Diffusion Model for Bimanual Manipulation

2026-03-07 · Zixuan Chen, Nga Teng Chan, Yiwen Hou, Chenrui Tie, Zixuan Liu, Haonan Chen, Junting Chen, Jieqi Shi, Yang Gao, Jing Huo, Lin Shao arxiv

Bimanual manipulation is a fundamental robotic skill that requires continuous and precise coordination between two arms. While imitation learning (IL) is the dominant paradigm for acquiring this capability, existing approaches, whether robot-centric or object-centric, often overlook the dynamic geometric relationship among the two arms and the manipulated object. This limitation frequently leads to inter-arm collisions, unstable grasps, and degraded performance in complex tasks. To address this, in this paper we explicitly models the Robot-Object Triadic Interaction (RoTri) representation in bimanual systems, by encoding the relative 6D poses between the two arms and the object to capture their spatial triadic relationship and establish continuous triangular geometric constraints. Building on this, we further introduce RoTri-Diff, a diffusion-based imitation learning framework that combines RoTri constraints with robot keyposes and object motion in a hierarchical diffusion process. This enables the generation of stable, coordinated trajectories and robust execution across different modes of bimanual manipulation. Extensive experiments show that our approach outperforms state-of-the-art baselines by 10.2% on 11 representative RLBench2 tasks and achieves stable performance on 4 challenging real-world bimanual tasks. Project website: https://rotri-diff.github.io/.

📄 PDF Abstract BibTeX arXiv:2603.07165

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Protriever: End-to-End Differentiable Protein Homology Search for Fitness Prediction

2025-06-10 · Ruben Weitzman, Peter Mørch Groth, Lood Van Niekerk, Aoi Otani 외

Retrieving homologous protein sequences is essential for a broad range of protein modeling tasks such as fitness prediction, protein design, structure modeling, and protein-protein interactions. Traditional workflows hav…

Protein DesignRetrieval

TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation

2026-05-07 · Hanyu Zhou, Chuanhao Ma, Gim Hee Lee arxiv

Vision-language-action (VLA) models perform well on training-seen robotic tasks but struggle to generalize to unseen scenes and objects. A key limitation lies in their implicit visual representations, which entangle obje…

ProTrix: Building Models for Planning and Reasoning over Tables with Sentence Context

2024-03-04 · Zirui Wu, Yansong Feng

Tables play a crucial role in conveying information in various domains. We propose a Plan-then-Reason framework to answer different types of user queries over tables with sentence context. The framework first plans the r…

In-Context LearningSentence

Pack It My Way: Triadic Human-Robot Collaboration for Personalized Autonomous Packing

2026-09-04 · Sandeep Chowdary Kotapati, Yanxin Gao, Tsung-Chi Lin arxiv

Personalized autonomous packing requires robots to account for resident preferences that cannot be inferred from scene geometry alone. Expert teleoperators can interpret these preferences and translate them into feasible…

Triadic Exploration and Exploration with Multiple Experts

2021-02-04 · Maximilian Felde, Gerd Stumme

Formal Concept Analysis (FCA) provides a method called attribute exploration which helps a domain expert discover structural dependencies in knowledge domains that can be represented by a formal context (a cross table of…

Attribute