paper-with-me

Papers

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments

2026-04-20 · Kangan Qian, ChuChu Xie, Yang Zhong, Jingrui Pang, Siwen Jiao, Sicong Jiang, Zilin Huang, Yunlong Wang, Kun Jiang, Mengmeng Yang, Hao Ye, Guanghao Zhang, Hangjun Ye, Guang Chen, Long Chen, Diange Yang arxiv

Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from complex environments. Current cloud pipelines rely on generic vision-language models (VLMs) that lack geometric reasoning and domain semantics due to their 2D image-text pretraining. To address this mismatch, we propose XEmbodied, a cloud-side foundation model that endows VLMs with intrinsic 3D geometric awareness and interaction with physical cues (e.g., occupancy grids, 3D boxes). Instead of treating geometry as auxiliary input, XEmbodied integrates geometric representations via a structured 3D Adapter and distills physical signals into context tokens using an Efficient Image-Embodied Adapter. Through progressive domain curriculum and reinforcement learning post-training, XEmbodied preserves general capabilities while demonstrating robust performance across 18 public benchmarks. It significantly improves spatial reasoning, traffic semantics, embodied affordance, and out-of-distribution generalization for large-scale scenario mining and embodied VQA.

📄 PDF Abstract BibTeX arXiv:2604.18484

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningSpatial Reasoning

Similar Papers 제목 키워드 기반

BED-SAM2: Boundary-Enhanced-Depth SAM2 via Monocular Geometric Priors

2026-05-24 · Tyler Rust, Dara McNally, Kyle O'Donnell, Colin Kelly 외 arxiv

Building upon the SAM2 vision foundation model for downstream segmentation, this study introduces Boundary Enhanced Depth (BED)-SAM2. The SAM2 Hiera encoder architecture is modified to directly encode monocular depth inf…

Object Detection

OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer

2025-11-13 · Haosong Peng, Hao Li, Yalun Dai, Yushi Lan 외 arxiv

General 3D foundation models have started to lead the trend of unifying diverse vision tasks, yet most assume RGB-only inputs and ignore readily available geometric cues (e.g., camera intrinsics, poses, and depth maps). …

Camera Pose EstimationDepth Estimation

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

2026-03-19 · Xianjin Wu, Dingkang Liang, Tianrui Feng, Kui Xia 외 arxiv

While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and physical dynamics. Existing solutions ty…

Scene UnderstandingSpatial ReasoningVideo Generation

CFSR: Geometry-Conditioned Shadow Removal via Physical Disentanglement

2026-04-20 · Pan Wang, Yihao Hu, Xiujin Liu, Hang Wang arxiv

Traditional shadow removal networks often treat image restoration as an unconstrained mapping, lacking the physical interpretability required to balance localized texture recovery with global illumination consistency. To…

Image RestorationShadow Removal

Vision Pretraining for Dense Spatial Perception

2026-07-06 · Zelin Fu, Bin Tan, Changjiang Sun, Shaohui Liu 외 arxiv

Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observations. Modern visual foundation models te…

Depth CompletionDepth Estimation