paper-with-me

Papers

A Map-free Deep Learning-based Framework for Gate-to-Gate Monocular Visual Navigation aboard Miniaturized Aerial Vehicles

2025-03-07 · Lorenzo Scarciglia, Antonio Paolillo, Daniele Palossi

Palm-sized autonomous nano-drones, i.e., sub-50g in weight, recently entered the drone racing scenario, where they are tasked to avoid obstacles and navigate as fast as possible through gates. However, in contrast with their bigger counterparts, i.e., kg-scale drones, nano-drones expose three orders of magnitude less onboard memory and compute power, demanding more efficient and lightweight vision-based pipelines to win the race. This work presents a map-free vision-based (using only a monocular camera) autonomous nano-drone that combines a real-time deep learning gate detection front-end with a classic yet elegant and effective visual servoing control back-end, only relying on onboard resources. Starting from two state-of-the-art tiny deep learning models, we adapt them for our specific task, and after a mixed simulator-real-world training, we integrate and deploy them aboard our nano-drone. Our best-performing pipeline costs of only 24M multiply-accumulate operations per frame, resulting in a closed-loop control performance of 30 Hz, while achieving a gate detection root mean square error of 1.4 pixels, on our ~20k real-world image dataset. In-field experiments highlight the capability of our nano-drone to successfully navigate through 15 gates in 4 min, never crashing and covering a total travel distance of ~100m, with a peak flight speed of 1.9 m/s. Finally, to stress the generalization capability of our system, we also test it in a never-seen-before environment, where it navigates through gates for more than 4 min.

📄 PDF Abstract BibTeX arXiv:2503.05251

Code (0)

등록된 구현이 없습니다.

Tasks

NavigateVisual Navigation

Methods 이 논문이 사용한 방법론

Travel 설명 없음
SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

ShapeGaussian: High-Fidelity 4D Human Reconstruction in Monocular Videos via Vision Priors

2026-02-05 · Zhenxiao Liang, Ning Zhang, Youbao Tang, Ruei-Sung Lin 외 arxiv

We introduce ShapeGaussian, a high-fidelity, template-free method for 4D human reconstruction from casual monocular videos. Generic reconstruction methods lacking robust vision priors, such as 4DGS, struggle to capture h…

Pose Estimation

Monocular Human Digitization via Implicit Re-projection Networks

2022-05-13 · Min-Gyu Park, Ju-Mi Kang, Je Woo Kim, Ju Hong Yoon

We present an approach to generating 3D human models from images. The key to our framework is that we predict double-sided orthographic depth maps and color images from a single perspective projected image. Our framework…

PromptMono: Cross Prompting Attention for Self-Supervised Monocular Depth Estimation in Challenging Environments

2025-01-23 · Changhao Wang, Guanwen Zhang, Zhengyun Cheng, Wei Zhou

Considerable efforts have been made to improve monocular depth estimation under ideal conditions. However, in challenging environments, monocular depth estimation still faces difficulties. In this paper, we introduce vis…

Depth EstimationMonocular Depth EstimationPrompt LearningSelf-Supervised Learning

Accelerating Transformer-Based Monocular SLAM via Geometric Utility Scoring

2026-04-09 · Xinmiao Xiong, Bangya Liu, Hao Wang, Dayou Li 외 arxiv

Geometric Foundation Models (GFMs) have recently advanced monocular SLAM by providing robust, calibration-free 3D priors. However, deploying these models on dense video streams introduces significant computational redund…

SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix

2024-06-29 · Peng Dai, Feitong Tan, Qiangeng Xu, David Futschik 외

Video generation models have demonstrated great capabilities of producing impressive monocular videos, however, the generation of 3D stereoscopic video remains under-explored. We propose a pose-free and training-free app…

DenoisingVideo GenerationVideo Inpainting