paper-with-me

Papers

Cross-Modal Benchmarking for Robotic Perception in Natural Environments

2026-06-10 · David Hall, Joshua Knights, Mark Cox, Peyman Moghadam arxiv

Natural environments present a complex challenge to robotics perception systems. Current models, particularly vision foundation models, are largely trained on structured, urban environments leading to weaknesses in their perception for field robotics tasks. We showcase the limitations of current models using our recently released WildCross benchmark, a new cross-modal benchmark for place recognition and metric depth estimation in large-scale natural environments. WildCross comprises over 476K sequential RGB frames with semi-dense depth and surface normal annotations, each aligned with accurate 6DoF pose and synchronized dense lidar submaps. In this work, we provide an expanded analysis of the benchmark results from the recent WildCross benchmark, with particular emphasis on expanded metric depth estimation experiments. Access to the code repository and dataset for this work can be found at https://csiro-robotics.github.io/WildCross.

📄 PDF Abstract BibTeX arXiv:2606.11563

Code (0)

등록된 구현이 없습니다.

Tasks

Depth Estimation

Similar Papers 제목 키워드 기반

WildCross: A Cross-Modal Large Scale Benchmark for Place Recognition and Metric Depth Estimation in Natural Environments

2026-03-02 · Joshua Knights, Joseph Reid, Kaushik Roy, David Hall 외 arxiv

Recent years have seen a significant increase in demand for robotic solutions in unstructured natural environments, alongside growing interest in bridging 2D and 3D scene understanding. However, existing robotics dataset…

Scene UnderstandingDepth Estimation

The Rosario Dataset v2: Multimodal Dataset for Agricultural Robotics

2025-08-29 · Nicolas Soncini, Javier Cremona, Erica Vidal, Maximiliano García 외 arxiv

We present a multi-modal dataset collected in a soybean crop field, comprising over two hours of recorded data from sensors such as stereo infrared camera, color camera, accelerometer, gyroscope, magnetometer, GNSS (Sing…

Large Language Models for Robotics: Opportunities, Challenges, and Perspectives

2024-01-09 · Jiaqi Wang, Zihao Wu, Yiwei Li, Hanqi Jiang 외

Large language models (LLMs) have undergone significant expansion and have been increasingly integrated across various domains. Notably, in the realm of robot task planning, LLMs harness their advanced reasoning and lang…

Robot Task PlanningTask Planning

TactEx: An Explainable Multimodal Robotic Interaction Framework for Human-Like Touch and Hardness Estimation

2026-02-21 · Felix Verstraete, Lan Wei, Wen Fan, Dandan Zhang arxiv

Accurate perception of object hardness is essential for safe and dexterous contact-rich robotic manipulation. Here, we present TactEx, an explainable multimodal robotic interaction framework that unifies vision, touch, a…

GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents

2026-03-25 · Yunzhe Wang, Runhui Xu, Kexin Zheng, Tianyi Zhang 외 arxiv

Multimodal LLMs are increasingly deployed as perceptual backbones for autonomous agents in 3D environments, from robotics to virtual worlds. These applications require agents to perceive rapid state changes, attribute ac…

Video Grounding