paper-with-me

홈 › Papers

OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies

2026-03-09 · Yunzhou Song, Long Le, Yong-Hyun Park, Jie Wang, Junyao Shi, Lingjie Liu, Jiatao Gu, Eric Eaton, Dinesh Jayaraman, Kostas Daniilidis arxiv

Vision-language-action(VLA) models have shown great promise as generalist policies for a large range of relatively simple tasks. However, they demonstrate limited performance on more complex tasks, such as those requiring complex spatial or semantic understanding, manipulation in clutter, or precise manipulation. We propose OMNIGUIDE, a flexible framework that improves VLA performance on such tasks by leveraging arbitrary sources of guidance, such as 3D foundation models, semantic-reasoning VLMs, and human pose models. We show how many kinds of guidance can be naturally expressed as differentiable energy functions with task-specific attractors and repellers located in 3D space, that influence the sampling of VLA actions. In this way, OMNIGUIDE enables guidance sources with complementary task-relevant strengths to improve a VLA model's performance on challenging tasks. Extensive experiments in both simulation and real-world environments, across diverse sources of guidance, demonstrate that OMNIGUIDE enhances the performance of state-of-the-art generalist policies (e.g., $π_{0.5}$, GR00T N1.6) significantly across success and safety rates. Critically, our unified framework matches or surpasses the performance of prior methods designed to incorporate specific sources of guidance into VLA policies. Project Page: $\href{https://omniguide.github.io/}{this \; url}$

📄 PDF Abstract BibTeX arXiv:2603.10052

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DGSolver: Diffusion Generalist Solver with Universal Posterior Sampling for Image Restoration

2025-04-30 · Hebaixu Wang, Jing Zhang, HaoNan Guo, Di Wang 외

Diffusion models have achieved remarkable progress in universal image restoration. While existing methods speed up inference by reducing sampling steps, substantial step intervals often introduce cumulative errors. Moreo…

Image RestorationNoise Estimation

Chameleon: A Data-Efficient Generalist for Dense Visual Prediction in the Wild

2024-04-29 · Donggyun Kim, Seongwoong Cho, Semin Kim, Chong Luo 외

Large language models have evolved data-efficient generalists, benefiting from the universal language interface and large-scale pre-training. However, constructing a data-efficient generalist for dense visual prediction …

Meta-Learning

Vision Foundation Models as Generalist Tokenizers for Image Generation

2026-05-18 · Anlin Zheng, Qi Han, Xin Wen, Chuofan Ma 외 arxiv

In this work, we explore the largely unexplored direction of building a generalist image tokenizer directly on top of a frozen vision foundation model (VFM). To build this tokenizer, we utilize a frozen VFM as the encode…

Self-Supervised LearningContrastive LearningImage Generation

Do as the Romans Do: Learning Universal Behaviors from Heterogeneous Agents

2026-06-16 · Caleb Chang, Davin Win Kyi, Natasha Jaques, Karen Leung arxiv

Humans often acquire new skills by observing others, since observed behaviors implicitly reveal how to act in an environment. However, observations drawn from a heterogeneous population introduce conflicting behavioral s…

Autonomous Driving

UniDexGrasp++: Improving Dexterous Grasping Policy Learning via Geometry-aware Curriculum and Iterative Generalist-Specialist Learning

2023-04-02 · ICCV 2023 1 · Weikang Wan, Haoran Geng, Yun Liu, Zikang Shan 외

We propose a novel, object-agnostic method for learning a universal policy for dexterous object grasping from realistic point cloud observations and proprioceptive information under a table-top setting, namely UniDexGras…

Object