paper-with-me

홈 › Papers

Beyond Isolation: A Unified Benchmark for General-Purpose Navigation

2026-05-10 · Samson Sun, Tianyi Yang, Tengyue Wang, Yikai Xue, Zhengjie Xu, Lingming Zhang, Qichen Zhang, Chao Liang, Zhipeng Zhang arxiv

The pursuit of general-purpose embodied agents is hindered by fragmented evaluation protocols that isolate navigation skills and fixate on specific robot morphologies, failing to reflect real-world scenarios where agents must orchestrate diverse behaviors across varying embodiments. To bridge this gap, we introduce OmniNavBench, a benchmark for cross-skill coordination and cross-embodiment generalization. OmniNavBench introduces three paradigm shifts: (1) Compositional Complexity. We propose composite instructions that interleave sub-tasks from 6 categories (PointNav, VLN, ObjectNav, SocialNav, Human Following and EQA), compelling agents to transition between exploration, interaction, and social compliance within a single episode. (2) Morphological Universality and Sensor Flexibility. We present a simulation platform that breaks the reliance on single-morphology evaluation, enabling generalization tests across humanoid, quadrupedal, and wheeled robots, with a modular sensor interface and 170 environments blending synthetic assets with real-world scans. (3) Demonstrations Quality. Moving beyond shortest-path algorithms, we curate 1779 expert trajectories via human teleoperation, capturing behavioral nuances such as exploratory glance and anticipatory avoidance. Extensive evaluations demonstrate that current methods, despite their claimed unified design, struggle with the complex, interleaved nature of general-purpose navigation. This exposes a critical disparity between existing capabilities and real-world deployment demands, underscoring OmniNavBench as a testbed for the next generation of generalist navigators. Dataset, code, and leaderboard are available at http://omninavbench.cloud-ip.cc.

📄 PDF Abstract BibTeX arXiv:2605.09441

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RapVerse: Coherent Vocals and Whole-Body Motions Generations from Text

2024-05-30 · Jiaben Chen, Xin Yan, Yihang Chen, Siyuan Cen 외

In this work, we introduce a challenging task for simultaneously generating 3D holistic body motions and singing vocals directly from textual lyrics inputs, advancing beyond existing works that typically address these tw…

Motion Generation

The Telephone Game: Evaluating Semantic Drift in Unified Models

2025-09-04 · Sabbir Mollah, Rohit Gupta, Sirnam Swetha, Qingyang Liu 외 arxiv

Employing a single, unified model (UM) for both visual understanding (image-to-text: I2T) and visual generation (text-to-image: T2I) has opened a new direction in Visual Language Model (VLM) research. While UMs can also …

RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark

2025-09-29 · Yang Shi, Yuhao Dong, Yue Ding, Yuran Wang 외 arxiv

The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question remains unanswered by existing benchmarks: …

Image Generation

Unified 3D Scene Understanding Through Physical World Modeling

2026-05-23 · Wanhee Lee, Klemen Kotar, Rahul Mysore Venkatesh, Jared Watrous 외 arxiv

Understanding 3D scenes requires flexible combinations of visual reasoning tasks, including depth estimation, novel view synthesis, and object manipulation, all of which are essential for perception and interaction. Exis…

Novel View SynthesisScene UnderstandingDepth EstimationVisual Reasoning

IE as Cache: Information Extraction Enhanced Agentic Reasoning

2026-04-16 · Hang Lv, Sheng Liang, Hongchao Gu, Wei Guo 외 arxiv

Information Extraction aims to distill structured, decision-relevant information from unstructured text, serving as a foundation for downstream understanding and reasoning. However, it is traditionally treated merely as …

Information Extraction