paper-with-me

홈 › Papers

MDE-AgriVLN: Agricultural Vision-and-Language Navigation with Monocular Depth Estimation

2025-12-03 · Xiaobei Zhao, Xingqi Lyu, Xin Chen, Xiang Li arxiv

Agricultural robots are serving as powerful assistants across a wide range of agricultural tasks, nevertheless, still heavily relying on manual operations or railway systems for movement. The AgriVLN method and the A2A benchmark pioneeringly extended Vision-and-Language Navigation (VLN) to the agricultural domain, enabling a robot to navigate to a target position following a natural language instruction. Unlike human binocular vision, most agricultural robots are only given a single camera for monocular vision, which results in limited spatial perception. To bridge this gap, we present the method of Agricultural Vision-and-Language Navigation with Monocular Depth Estimation (MDE-AgriVLN), in which we propose the MDE module generating depth features from RGB images, to assist the decision-maker on multimodal reasoning. When evaluated on the A2A benchmark, our MDE-AgriVLN method successfully increases Success Rate from 0.23 to 0.32 and decreases Navigation Error from 4.43m to 4.08m, demonstrating the state-of-the-art performance in the agricultural VLN domain. Code: https://github.com/AlexTraveling/MDE-AgriVLN.

📄 PDF Abstract BibTeX arXiv:2512.03958

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth EstimationMultimodal Reasoning

Similar Papers 제목 키워드 기반

SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation

2025-10-16 · Xiaobei Zhao, Xingqi Lyu, Xiang Li arxiv

Agricultural robots are emerging as powerful assistants across a wide range of agricultural tasks, nevertheless, they are still heavily relying on manual operations or fixed railways for movement. The A2A benchmark and t…

3D Reconstruction

AgriVLN: Vision-and-Language Navigation for Agricultural Robots

2025-08-10 · Xiaobei Zhao, Xingqi Lyu, Xiang Li arxiv

Agricultural robots have emerged as powerful members in agricultural tasks, nevertheless, still heavily rely on manual operation or untransportable railway for movement, resulting in limited mobility and poor adaptabilit…

TEA-AgriVLN: Traversability Estimation Alarm for Agricultural Vision-and-Language Navigation

2026-07-30 · Xiaobei Zhao, Xingqi Lyu, Xin Chen, Xiang Li arxiv

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow a natural language instruction, predicting a sequence of low-level actions to navigate a robot from a starting point to a tar…

T-araVLN: Translator for Agricultural Robotic Agents on Vision-and-Language Navigation

2025-09-08 · Xiaobei Zhao, Xingqi Lyu, Xin Chen, Xiang Li arxiv

Agricultural robotic agents have been becoming useful helpers in a wide range of agricultural tasks. However, they still heavily rely on manual operations or fixed railways for movement. To address this limitation, the A…

IMAC-AgriVLN: Can Agricultural Vision-and-Language Navigation Agents be Aware of Instruction Mistakes?

2026-06-01 · Xiaobei Zhao, Xingqi Lyu, Xin Chen, Xiang Li arxiv

Agricultural robots are serving as powerful assistants across a wide range of agricultural tasks, nevertheless, still heavily relying on manual operations or railway systems for movement. The AgriVLN method and the A2A b…