paper-with-me

Papers

Revisiting Birds Eye View Perception Models with Frozen Foundation Models: DINOv2 and Metric3Dv2

2025-01-14 · Seamie Hayes, Ganesh Sistu, Ciarán Eising

Birds Eye View perception models require extensive data to perform and generalize effectively. While traditional datasets often provide abundant driving scenes from diverse locations, this is not always the case. It is crucial to maximize the utility of the available training data. With the advent of large foundation models such as DINOv2 and Metric3Dv2, a pertinent question arises: can these models be integrated into existing model architectures to not only reduce the required training data but surpass the performance of current models? We choose two model architectures in the vehicle segmentation domain to alter: Lift-Splat-Shoot, and Simple-BEV. For Lift-Splat-Shoot, we explore the implementation of frozen DINOv2 for feature extraction and Metric3Dv2 for depth estimation, where we greatly exceed the baseline results by 7.4 IoU while utilizing only half the training data and iterations. Furthermore, we introduce an innovative application of Metric3Dv2's depth information as a PseudoLiDAR point cloud incorporated into the Simple-BEV architecture, replacing traditional LiDAR. This integration results in a +3 IoU improvement compared to the Camera-only model.

📄 PDF Abstract BibTeX arXiv:2501.08118

Code (0)

등록된 구현이 없습니다.

Tasks

Depth Estimation

Similar Papers 제목 키워드 기반

AsyncMDE: Real-Time Monocular Depth Estimation via Asynchronous Spatial Memory

2026-03-11 · Lianjie Ma, Yuquan Li, Bingzheng Jiang, Ziming Zhong 외 arxiv

Foundation-model-based monocular depth estimation offers a viable alternative to active sensors for robot perception, yet its computational cost often prohibits deployment on edge platforms. Existing methods perform inde…

Monocular Depth Estimation

FishRoPE: Projective Rotary Position Embeddings for Omnidirectional Visual Perception

2026-04-12 · Rahul Ahuja, Mudit Jain, Bala Murali Manoghar Sai Sudhakar, Venkatraman Narayanan 외 arxiv

Vision foundation models (VFMs) and Bird's Eye View (BEV) representation have advanced visual perception substantially, yet their internal spatial representations assume the rectilinear geometry of pinhole cameras. Fishe…

Autonomous VehiclesBEV Segmentation

Foundation Models for Bioacoustics -- a Comparative Review

2025-08-02 · Raphael Schwinger, Paria Vali Zadeh, Lukas Rauch, Mats Kurz 외 arxiv

Automated bioacoustic analysis is essential for biodiversity monitoring and conservation, requiring advanced deep learning models that can adapt to diverse bioacoustic tasks. This article presents a comprehensive review …

Self-Supervised LearningRepresentation Learning

A review of mass concentrations of Bramblings Fringilla montifringilla: implications for assessment of large numbers of birds

2020-10-24

Mass concentrations of birds, or lack of such, is a phenomenon of great ecological and domestic significance. Apart from being and indicator for e.g. food availability, ecological change and population size, it is also a…

Multi-view Tracking, Re-ID, and Social Network Analysis of a Flock of Visually Similar Birds in an Outdoor Aviary

2022-12-01 · Shiting Xiao, Yufu Wang, Ammon Perkes, Bernd Pfrommer 외

The ability to capture detailed interactions among individuals in a social group is foundational to our study of animal behavior and neuroscience. Recent advances in deep learning and computer vision are driving rapid pr…

3D Reconstruction