paper-with-me

Papers

Learning Structural Latent Points for Efficient Visual Representations in Robotic Manipulation

2026-05-20 · Yicheng Jiang, Jiaxu Wang, Junhao He, Zesen Gan, Junhao Li, Qiang Zhang, Jingkai Sun, Jiahang Cao, Mingyuan Sun, Xiangyu Yue, Qiming Shao arxiv

Current 3D-aware pretraining methods for embodied perception and manipulation are largely built on differentiable rendering frameworks, producing either fully implicit neural fields or fully explicit geometric primitives. Implicit representations, while expressive, lack explicit structural cues, whereas explicit ones preserve geometry but suffer from resolution limits and weak generalization. To address these limitations, we propose a novel pretraining framework that learns a hybrid representation-structural latent points. Specifically, we insert a point-wise latent variational autoencoder into the latent space of a point-cloud autoencoder, jointly regularizing point-wise features and coordinates toward a Gaussian prior. The resulting compact latent preserves coarse structural tendencies, which do not encode precise geometry but capture richer rough shape and semantic information, effectively combining the expressiveness of implicit representations with the structural priors of explicit ones. In addition, informed by shared design choices in prior work, we develop a streamlined, efficient 3DGS-based rendering pipeline that is deliberately kept lightweight, improving efficiency while leaving greater representational capacity to the front-end latent module. Extensive evaluations on RLBench, ManiSkill2, and a real-robot platform demonstrate consistent gains in task success, sample efficiency, and robustness to viewpoint and scene variations over strong baselines. Ablation studies further confirm that each component of our framework is critical to overall performance.

📄 PDF Abstract BibTeX arXiv:2605.21258

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unsupervised Learning of Visual 3D Keypoints for Control

2021-06-14 · Boyuan Chen, Pieter Abbeel, Deepak Pathak

Learning sensorimotor control policies from high-dimensional images crucially relies on the quality of the underlying visual representations. Prior works show that structured latent space such as visual keypoints often o…

Learning to Act Robustly with View-Invariant Latent Actions

2026-01-06 · Youngjoon Jeong, Junha Chun, Taesup Kim arxiv

Vision-based robotic policies often struggle with even minor viewpoint changes, underscoring the need for view-invariant visual representations. This challenge becomes more pronounced in real-world settings, where viewpo…

LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model

2026-01-08 · Zhuoyang Liu, Jiaming Liu, Hao Chen, Jiale Yu 외 arxiv

Vision-Language-Action (VLA) models have recently shown strong generalization, with some approaches seeking to explicitly generate linguistic reasoning traces or predict future observations prior to execution. However, e…

Embody4D: A Generalist Data Engine for Embodied 4D World Modeling

2026-05-03 · Peiyan Tu, Hanxin Zhu, Jingwen Sun, Shaojie Ren 외 arxiv

Embodied agents require robust and comprehensive 3D spatiotemporal representations to support spatial reasoning, manipulation understanding, and downstream decision making. However, existing robot data are typically capt…

Spatial ReasoningDecision Making

DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation

2026-08-22 · Fenghao Lei, Zhixiong Huang, Long Yang, Jiabao Chen 외 arxiv

World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generation is an unnecessary intermediate object…

Video Generation