paper-with-me

Papers

Towards Transferable Multi-modal Perception Representation Learning for Autonomy: NeRF-Supervised Masked AutoEncoder

2023-11-23 · Xiaohao Xu

This work proposes a unified self-supervised pre-training framework for transferable multi-modal perception representation learning via masked multi-modal reconstruction in Neural Radiance Field (NeRF), namely NeRF-Supervised Masked AutoEncoder (NS-MAE). Specifically, conditioned on certain view directions and locations, multi-modal embeddings extracted from corrupted multi-modal input signals, i.e., Lidar point clouds and images, are rendered into projected multi-modal feature maps via neural rendering. Then, original multi-modal signals serve as reconstruction targets for the rendered multi-modal feature maps to enable self-supervised representation learning. Extensive experiments show that the representation learned via NS-MAE shows promising transferability for diverse multi-modal and single-modal (camera-only and Lidar-only) perception models on diverse 3D perception downstream tasks (3D object detection and BEV map segmentation) with diverse amounts of fine-tuning labeled data. Moreover, we empirically find that NS-MAE enjoys the synergy of both the mechanism of masked autoencoder and neural radiance field. We hope this study can inspire exploration of more general multi-modal representation learning for autonomous agents.

📄 PDF Abstract BibTeX arXiv:2311.13750

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionNeRFNeural Renderingobject-detectionObject DetectionRepresentation Learning

Similar Papers 제목 키워드 기반

TinyBEV: Cross Modal Knowledge Distillation for Efficient Multi Task Bird's Eye View Perception and Planning

2025-09-22 · Reeshad Khan, John Gauch arxiv

We present TinyBEV, a unified, camera only Bird's Eye View (BEV) framework that distills the full-stack capabilities of a large planning-oriented teacher (UniAD [19]) into a compact, real-time student model. Unlike prior…

Knowledge DistillationMotion Forecasting

UniPilot: Enabling GPS-Denied Autonomy Across Embodiments

2025-09-15 · Mihir Kulkarni, Mihir Dharmadhikari, Nikhil Khedekar, Morten Nissov 외 arxiv

This paper presents UniPilot, a compact hardware-software autonomy payload that can be integrated across diverse robot embodiments to enable autonomous operation in GPS-denied environments. The system integrates a multi-…

TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation

2025-05-15 · Manthan Patel, Fan Yang, Yuheng Qiu, Cesar Cadena 외

We present TartanGround, a large-scale, multi-modal dataset to advance the perception and autonomy of ground robots operating in diverse environments. This dataset, collected in various photorealistic simulation environm…

Optical Flow Estimation

SeePerSea: Multi-modal Perception Dataset of In-water Objects for Autonomous Surface Vehicles

2024-04-29 · Mingi Jeong, Arihant Chadda, Ziang Ren, Luyang Zhao 외

This paper introduces the first publicly accessible labeled multi-modal perception dataset for autonomous maritime navigation, focusing on in-water obstacles within the aquatic environment to enhance situational awarenes…

object-detectionObject Detection

Adaptive Forensic Feature Refinement via Intrinsic Importance Perception

2026-04-18 · Jiazhen Yang, Junjun Zheng, Kejia Chen, Xiangheng Kong 외 arxiv

With the rapid development of generative models and multimodal content editing technologies, the key challenge faced by synthetic image detection (SID) lies in cross-distribution generalization to unknown generation sour…