paper-with-me

홈 › Papers

FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection

2026-03-09 · Anqi Joyce Yang, James Tu, Nikita Dvornik, Enxu Li, Raquel Urtasun arxiv

In order to navigate complex traffic environments, self-driving vehicles must recognize many semantic classes pertaining to vulnerable road users or traffic control devices. However, many safety-critical objects (e.g., construction worker) appear infrequently in nominal traffic conditions, leading to a severe shortage of training examples from driving data alone. Recent vision foundation models, which are trained on a large corpus of data, can serve as a good source of external prior knowledge to improve generalization. We propose FOMO-3D, the first multi-modal 3D detector to leverage vision foundation models for long-tailed 3D detection. Specifically, FOMO-3D exploits rich semantic and depth priors from OWLv2 and Metric3Dv2 within a two-stage detection paradigm that first generates proposals with a LiDAR-based branch and a novel camera-based branch, and refines them with attention especially to image features from OWL. Evaluations on real-world driving data show that using rich priors from vision foundation models with careful multi-modal fusion designs leads to large gains for long-tailed 3D detection. Project website is at https://waabi.ai/fomo3d/.

📄 PDF Abstract BibTeX arXiv:2603.08611

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object Detection

Similar Papers 제목 키워드 기반

FoMo4Wheat: Toward reliable crop vision foundation models with globally curated data

2025-09-08 · Bing Han, Chen Zhu, Dong Han, Rui Yu 외 arxiv

Vision-driven field monitoring is central to digital agriculture, yet models built on general-domain pretrained backbones often fail to generalize across tasks, owing to the interaction of fine, variable canopy structure…

Towards Brain MRI Foundation Models for the Clinic: Findings from the FOMO25 Challenge

2026-04-13 · Asbjørn Munk, Stefano Cerri, Vardan Nersesjan, Christian Hedeager Krag 외 arxiv

Clinical deployment of automated brain MRI analysis faces a fundamental challenge: clinical data is heterogeneous and noisy, and high-quality labels are prohibitively costly to obtain. Self-supervised learning (SSL) can …

Self-Supervised Learning

FoMo-Bench: a multi-modal, multi-scale and multi-task Forest Monitoring Benchmark for remote sensing foundation models

2023-12-15 · Nikolaos Ioannis Bountos, Arthur Ouaknine, David Rolnick

Forests are an essential part of Earth's ecosystems and natural systems, as well as providing services on which humanity depends, yet they are rapidly changing as a result of land use decisions and climate change. Unders…

object-detectionObject Detection

FoMo: A Foundation Model for Mobile Traffic Forecasting with Diffusion Model

2024-10-20 · Haoye Chai, Xiaoqian Qi, Shiyuan Zhang, Yong Li

Mobile traffic forecasting allows operators to anticipate network dynamics and performance in advance, offering substantial potential for enhancing service quality and improving user experience. However, existing models …

Contrastive LearningFew-Shot LearningmodelScheduling+1

Open World Object Detection in the Era of Foundation Models

2023-12-10 · Orr Zohar, Alejandro Lozano, Shelly Goel, Serena Yeung 외

Object detection is integral to a bevy of real-world applications, from robotics to medical image analysis. To be used reliably in such applications, models must be capable of handling unexpected - or novel - objects. Th…

Medical Image AnalysisObjectobject-detectionObject Detection+1