paper-with-me

Papers

SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation

2025-11-12 · Hao Shi, Bin Xie, Yingfei Liu, Yang Yue, Tiancai Wang, Haoqiang Fan, Xiangyu Zhang, Gao Huang arxiv

Robotic manipulation requires precise spatial understanding to interact with objects in the real world. Point-based methods suffer from sparse sampling, leading to the loss of fine-grained semantics. Image-based methods typically feed RGB and depth into 2D backbones pre-trained on 3D auxiliary tasks, but their entangled semantics and geometry are sensitive to inherent depth noise in real-world that disrupts semantic understanding. Moreover, these methods focus on high-level geometry while overlooking low-level spatial cues essential for precise interaction. We propose SpatialActor, a disentangled framework for robust robotic manipulation that explicitly decouples semantics and geometry. The Semantic-guided Geometric Module adaptively fuses two complementary geometry from noisy depth and semantic-guided expert priors. Also, a Spatial Transformer leverages low-level spatial cues for accurate 2D-3D mapping and enables interaction among spatial features. We evaluate SpatialActor on multiple simulation and real-world scenarios across 50+ tasks. It achieves state-of-the-art performance with 87.4% on RLBench and improves by 13.9% to 19.4% under varying noisy conditions, showing strong robustness. Moreover, it significantly enhances few-shot generalization to new tasks and maintains robustness under various spatial perturbations. Project Page: https://shihao1895.github.io/SpatialActor

📄 PDF Abstract BibTeX arXiv:2511.09555

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ManiVID-3D: Generalizable View-Invariant Reinforcement Learning for Robotic Manipulation via Disentangled 3D Representations

2025-09-14 · Zheng Li, Pei Qu, Yufei Jia, Shihui Zhou 외 arxiv

Deploying visual reinforcement learning (RL) policies in real-world manipulation is often hindered by camera viewpoint changes. A policy trained from a fixed front-facing camera may fail when the camera is shifted -- an …

Reinforcement Learning

Disentangled Latent Transformer for Interpretable Monocular Height Estimation

2022-01-17 · Zhitong Xiong, Sining Chen, Yilei Shi, Xiao Xiang Zhu

Monocular height estimation (MHE) from remote sensing imagery has high potential in generating 3D city models efficiently for a quick response to natural disasters. Most existing works pursue higher performance. However,…

Semantic SegmentationUnsupervised Semantic Segmentation

Exploring Disentanglement with Multilingual and Monolingual VQ-VAE

2021-05-04 · Jennifer Williams, Jason Fong, Erica Cooper, Junichi Yamagishi

This work examines the content and usefulness of disentangled phone and speaker representations from two separately trained VQ-VAE systems: one trained on multilingual data and another trained on monolingual data. We exp…

Disentanglement

Independence Promoted Graph Disentangled Networks

2019-11-26 · Yanbei Liu, Xiao Wang, Shu Wu, Zhitao Xiao

We address the problem of disentangled representation learning with independent latent factors in graph convolutional networks (GCNs). The current methods usually learn node representation by describing its neighborhood …

ClusteringGraph ClassificationGraph ClusteringRepresentation Learning

DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models

2025-09-17 · Kevin Wilkinghoff, Zheng-Hua Tan arxiv

Reasoning about spatial audio with large language models requires a spatial audio encoder as an acoustic front-end to obtain audio embeddings for further processing. Such an encoder needs to capture all information requi…