paper-with-me

Papers

Building Explicit World Model for Zero-Shot Open-World Object Manipulation

2026-03-14 · Xiaotong Li, Gang Chen, Javier Alonso-Mora arxiv

Open-world object manipulation remains a fundamental challenge in robotics. While Vision-Language-Action (VLA) models have demonstrated promising results, they rely heavily on large-scale robot action demonstrations, which are costly to collect and can hinder out-of-distribution generalization. In this paper, we propose an explicit-world-model-based framework for open-world manipulation that achieves zero-shot generalization by constructing a physically grounded digital twin of the environment. The framework integrates open-set perception, digital-twin reconstruction, sampling and evaluation of interaction strategies. By constructing a digital twin of the environment, our approach efficiently explores and evaluates manipulation strategies in physic-enabled simulator and reliably deploys the chosen strategy to the real world. Experimentally, the proposed framework is able to perform multiple open-set manipulation tasks without any task-specific action demonstrations, proving strong zero-shot generalization on both the task and object levels. Project Page: https://bojack-bj.github.io/projects/thesis/

📄 PDF Abstract BibTeX arXiv:2603.13825

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot Generalization

Similar Papers 제목 키워드 기반

MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation

2025-11-13 · Xun Huang, Shijia Zhao, Yunxiang Wang, Xin Lu 외 arxiv

Embodied navigation is a fundamental capability for robotic agents operating. Real-world deployment requires open vocabulary generalization and low training overhead, motivating zero-shot methods rather than task-specifi…

Open-World Objectness Modeling Unifies Novel Object Detection

2025-01-01 · CVPR 2025 1 · Shan Zhang, Yao Ni, Jinhao Du, Yuan Xue 외

The challenge in open-world object detection, similarly to few- and zero-shot learning, is to generalize beyond the class distribution of the training data. In this paper, we propose a general class-agnostic objectne…

Novel Object Detectionobject-detectionObject DetectionOpen-vocabulary object detection+3

GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models

2025-05-30 · Gilles Quentin Hacheme, Girmaw Abebe Tadesse, Caleb Robinson, Akram Zaytar 외

Classifying geospatial imagery remains a major bottleneck for applications such as disaster response and land-use monitoring-particularly in regions where annotated data is scarce or unavailable. Existing tools (e.g., RS…

ClassificationDisaster Responseimage-classificationImage Classification+6

OpenFMNav: Towards Open-Set Zero-Shot Object Navigation via Vision-Language Foundation Models

2024-02-16 · Yuxuan Kuang, Hai Lin, Meng Jiang

Object navigation (ObjectNav) requires an agent to navigate through unseen environments to find queried objects. Many previous methods attempted to solve this task by relying on supervised or reinforcement learning, wher…

Common Sense ReasoningNavigate

Embracing Diversity: Interpretable Zero-shot classification beyond one vector per class

2024-04-25 · Mazda Moayeri, Michael Rabbat, Mark Ibrahim, Diane Bouchacourt

Vision-language models enable open-world classification of objects without the need for any retraining. While this zero-shot paradigm marks a significant advance, even today's best models exhibit skewed performance when …

Diversityzero-shot-classificationZero-Shot Learning