paper-with-me

홈 › Papers

See Selectively, Act Adaptively: Dual-Level Structural Decomposition for Bimanual Robot Manipulation

2026-06-11 · Yoon-Ji Choi, Young-Chae Son, Soo-Chul Lim arxiv

In bimanual robotic manipulation, task-relevant visual information varies with the task stage and context, while the interaction of the two arms shifts between independent and coordinated modes, making policy learning challenging. However, existing monolithic Vision-Language-Action (VLA) policies process diverse visual inputs and interaction patterns through a single shared representation and action generation pathway, often failing to separately account for visual relevance and bimanual interaction structure. To address this issue, we propose a bimanual manipulation VLA framework based on Dual-Level Structural Decomposition. The View-Selective Visual Router dynamically adjusts wrist-view contributions to emphasize relevant visual cues, while the Interaction-Aware Action Mixture-of-Experts (MoE) decomposes action generation into coordinated and arm-wise pathways to adapt to varying bimanual interaction modes. We evaluate the proposed method on six simulated bimanual manipulation tasks in RoboTwin 2.0 and three long-horizon real-world tasks. Our model improves the overall average success rate over a monolithic baseline by 27.7% in simulation and 43.3% in real-world evaluation, while consistently outperforming single-module variants across both settings. These results demonstrate that jointly considering selective visual processing and explicit decomposition of bimanual interaction structures provides an effective inductive bias for robust bimanual manipulation.

📄 PDF Abstract BibTeX arXiv:2606.13279

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

Structural and Statistical Texture Knowledge Distillation for Semantic Segmentation

2023-05-06 · CVPR 2022 1 · Deyi Ji, Haoran Wang, Mingyuan Tao, Jianqiang Huang 외

Existing knowledge distillation works for semantic segmentation mainly focus on transferring high-level contextual knowledge from teacher to student. However, low-level texture knowledge is also of vital importance for c…

Knowledge DistillationQuantizationSemantic Segmentation

SSDA: Bridging Spectral and Structural Gaps via Dual Adaptation for Vision-Based Time Series Forecasting

2026-05-10 · Mingrui Zhang, Hanchen Yang, Wengen Li, Xudong Jiang 외 arxiv

Large vision models (LVMs) have recently proven to be surprisingly effective time series forecasters, simply by rendering temporal data as images. This success, how ever, rests on a largely unexamined premise: the render…

Time Series ForecastingTemporal Sequences

GADPN: Graph Adaptive Denoising and Perturbation Networks via Singular Value Decomposition

2026-01-13 · Hao Deng, Bo Liu arxiv

While Graph Neural Networks (GNNs) excel on graph-structured data, their performance is fundamentally limited by the quality of the observed graph, which often contains noise, missing links, or structural properties misa…

Graph structure learning

Variational Degeneration to Structural Refinement: A Unified Framework for Superimposed Image Decomposition

2023-01-01 · ICCV 2023 1 · Wenyu Li, Yan Xu, Yang Yang, Haoran Ji 외

Decomposing a single mixed image into individual image layers is the common crux of a classical category of tasks in image restoration. Several unified frameworks have been proposed that can handle different types of…

Image RestorationImage Shadow RemovalRain RemovalReflection Removal+1

Residual-SwinCA-Net: A Channel-Aware Integrated Residual CNN-Swin Transformer for Malignant Lesion Segmentation in BUSI

2025-12-09 · Saeeda Naz, Saddam Hussain Khan arxiv

A novel deep hybrid Residual-SwinCA-Net segmentation framework is proposed in the study for addressing such challenges by extracting locally correlated and robust features, incorporating residual CNN modules. Furthermore…

Lesion Segmentation