paper-with-me

Papers

Inductive Generalization for Robotic Manipulation

2026-06-19 · Annabella Macaluso, Haochen Zhang, Ishaan Masilamony, Yingshan Chang, Yonatan Bisk arxiv

Understanding the generalization capabilities of visuomotor policies is essential in the development of capable robotic agents. Generalizable models learn structures that transfer across domains. However, in practice, visuomotor policies test performance by interpolation on known distributions using unstructured domain shifts (e.g. lighting, clutter, diverse objects). We argue that to measure generalization capabilities we must instead test the inductive capacity of policies on progressively harder, out-of-distribution task variants. We call this inductive generalization, drawing directly on how axis-based evaluation has revealed inherent generalization limitations in language models (e.g. sequence length, counting) arXiv:2502.00197 . We provide a reusable and formal evaluation protocol for measuring inductive generalization in any manipulation policy, and establish baselines showing that existing paradigms fail this test; e.g. SoTA Vision-Language-Action models and find that policies that appear to generalize to prior domain shifts (distractors, etc) fail inductive generalization tests. These results expose a class of learning challenges orthogonal to those addressed by data and model scaling in robot learning, yet are imperative to solve in order to realize general purpose robots.

📄 PDF Abstract BibTeX arXiv:2606.20999

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Object-Centric Representations Improve Policy Generalization in Robot Manipulation

2025-05-16 · Alexandre Chapin, Bruno Machado, Emmanuel Dellandrea, Liming Chen

Visual representations are central to the learning and generalization capabilities of robotic manipulation policies. While existing methods rely on global or dense features, such representations often entangle task-relev…

Optical Character Recognition (OCR)Robot Manipulation

Morphologically Equivariant Flow Matching for Bimanual Mobile Manipulation

2026-05-12 · Max Siebenborn, Daniel Ordoñez Apraez, Sophie Lueth, Giulio Turrisi 외 arxiv

Mobile manipulation requires coordinated control of high-dimensional, bimanual robots. Imitation learning methods have been broadly used to solve these robotic tasks, yet typically ignore the bilateral morphological symm…

Zero-shot Generalization

Disentangled Object-Centric Image Representation for Robotic Manipulation

2025-03-14 · David Emukpere, Romain Deffayet, Bingbing Wu, Romain Brégier 외

Learning robotic manipulation skills from vision is a promising approach for developing robotics applications that can generalize broadly to real-world scenarios. As such, many approaches to enable this vision have been …

Object

FUNCanon: Learning Pose-Aware Action Primitives via Functional Object Canonicalization for Generalizable Robotic Manipulation

2025-09-23 · Hongli Xu, Lei Zhang, Xiaoyue Hu, Boyang Zhong 외 arxiv

General-purpose robotic skills from end-to-end demonstrations often leads to task-specific policies that fail to generalize beyond the training distribution. Therefore, we introduce FunCanon, a framework that converts lo…

$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills

2026-04-27 · Siyao Xiao, Yuhong Zhang, Zhifang Liu, Zihan Gao 외 arxiv

Current Vision-Language-Action (VLA) models predominantly rely on end-to-end fine-tuning. While effective, this paradigm compromises the inherent generalization capabilities of Vision-Language Models (VLMs) and incurs ca…