paper-with-me

홈 › Papers

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

2026-07-21 · Ziqin Wang, Hao Li, Weijun Wang, Junhao Cai, Jia Zeng, Yilun Chen, Jiangmiao Pang, Si Liu arxiv

Existing robot datasets remain expensive to curate, embodiment-specific, and insufficiently annotated with the fine-grained structure required for generalizable reasoning, execution, or long-horizon environment dynamics simulation. Building on our prior work, RoboInter1.0, we present RoboInter1.5, an extended and holistic suite of intermediate representations for both robotic manipulation and embodied world modeling. RoboInter1.5 provides a unified resource of data, benchmarks, and models centered on dense manipulation-oriented intermediate representations. Specifically, RoboInter-Data contains over 230k manipulation episodes across 571 scenes with dense per-frame annotations covering more than ten types of intermediate representations, including subtasks, primitive skills, object and gripper grounding, segmentation, affordance, grasp poses, contact points, motion traces, etc. Built upon these annotations, RoboInter-VQA introduces spatial and temporal embodied VQA tasks to benchmark and improve the intermediate-representation reasoning capabilities of our RoboInter-VLM. RoboInter-VLA further studies how such representations benefit action execution through implicit, explicit, and modular plan-then-execute paradigms. To better model the physical world, we further introduce RoboInter-World, which leverages intermediate representations as structured conditioning signals for controllable prediction of future world states. Extensive evaluations demonstrate that RoboInter1.5 provides a unified spatiotemporal scaffolding for intermediate representations. Rather than treating intermediate representations merely as interpretable signals, RoboInter1.5 conceptualizes them as a bidirectional interface that both regularizes low-level action spaces and constrains the latent rollouts of open-world physical simulators.

📄 PDF Abstract BibTeX arXiv:2607.18709

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation

2026-02-10 · Hao Li, Ziqin Wang, Zi-han Ding, Shuai Yang 외 arxiv

Advances in large vision-language models (VLMs) have stimulated growing interest in vision-language-action (VLA) systems for robot manipulation. However, existing manipulation datasets remain costly to curate, highly emb…

Robot Manipulation

EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI

2023-12-26 · CVPR 2024 1 · Tai Wang, Xiaohan Mao, Chenming Zhu, Runsen Xu 외

In the realm of computer vision and robotics, embodied agents are expected to explore their environment and carry out human instructions. This necessitates the ability to fully understand 3D scenes given their first-pers…

Scene Understanding

LangSuitE: Planning, Controlling and Interacting with Large Language Models in Embodied Text Environments

2024-06-24 · Zixia Jia, Mengmeng Wang, Baichen Tong, Song-Chun Zhu 외

Recent advances in Large Language Models (LLMs) have shown inspiring achievements in constructing autonomous agents that rely on language descriptions as inputs. However, it remains unclear how well LLMs can function as …

World Knowledge

Holistic Sentence Embeddings for Better Out-of-Distribution Detection

2022-10-14 · Sishuo Chen, Xiaohan Bi, Rundong Gao, Xu sun

Detecting out-of-distribution (OOD) instances is significant for the safe deployment of NLP models. Among recent textual OOD detection works based on pretrained language models (PLMs), distance-based methods have shown s…

AvgOut-of-Distribution DetectionOut of Distribution (OOD) DetectionSentence+3

The Semantic Lifecycle in Embodied AI: Acquisition, Representation and Storage via Foundation Models

2026-01-12 · Shuai Chen, Hao Chen, Yuanchen Bei, Tianyang Zhao 외 arxiv

Semantic information in embodied AI is inherently multi-source and multi-stage, making it challenging to fully leverage for achieving stable perception-to-action loops in real-world environments. Early studies have combi…

Domain Generalization