paper-with-me

Papers

AugMapNet: Improving Spatial Latent Structure via BEV Grid Augmentation for Enhanced Vectorized Online HD Map Construction

2025-03-17 · Thomas Monninger, Md Zafar Anwar, Stanislaw Antol, Steffen Staab, Sihao Ding

Autonomous driving requires an understanding of the infrastructure elements, such as lanes and crosswalks. To navigate safely, this understanding must be derived from sensor data in real-time and needs to be represented in vectorized form. Learned Bird's-Eye View (BEV) encoders are commonly used to combine a set of camera images from multiple views into one joint latent BEV grid. Traditionally, from this latent space, an intermediate raster map is predicted, providing dense spatial supervision but requiring post-processing into the desired vectorized form. More recent models directly derive infrastructure elements as polylines using vectorized map decoders, providing instance-level information. Our approach, Augmentation Map Network (AugMapNet), proposes latent BEV grid augmentation, a novel technique that significantly enhances the latent BEV representation. AugMapNet combines vector decoding and dense spatial supervision more effectively than existing architectures while remaining as straightforward to integrate and as generic as auxiliary supervision. Experiments on nuScenes and Argoverse2 datasets demonstrate significant improvements in vectorized map prediction performance up to 13.3% over the StreamMapNet baseline on 60m range and greater improvements on larger ranges. We confirm transferability by applying our method to another baseline and find similar improvements. A detailed analysis of the latent BEV grid confirms a more structured latent space of AugMapNet and shows the value of our novel concept beyond pure performance improvement. The code will be released soon.

📄 PDF Abstract BibTeX arXiv:2503.13430

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingNavigate

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Generative Video Compression with One-Dimensional Latent Representation

2026-03-16 · Zihan Zheng, Zhaoyang Jia, Naifu Xue, Jiahao Li 외 arxiv

Recent advancements in generative video codec (GVC) typically encode video into a 2D latent grid and employ high-capacity generative decoders for reconstruction. However, this paradigm still leaves two key challenges in …

ALTO: Alternating Latent Topologies for Implicit 3D Reconstruction

2022-12-08 · CVPR 2023 1 · Zhen Wang, Shijie Zhou, Jeong Joon Park, Despoina Paschalidou 외

This work introduces alternating latent topologies (ALTO) for high-fidelity reconstruction of implicit 3D surfaces from noisy point clouds. Previous work identifies that the spatial arrangement of latent encodings is imp…

3D ReconstructionLicense Plate Recognition

A Spatial Model for Extracting and Visualizing Latent Discourse Structure in Text

2018-07-01 · ACL 2018 7 · Shashank Srivastava, Nebojsa Jojic

We present a generative probabilistic model of documents as sequences of sentences, and show that inference in it can lead to extraction of long-range latent discourse structure from a collection of documents. The approa…

Information RetrievalReading ComprehensionSemantic SimilaritySemantic Textual Similarity+3

Non-invasive Neural Decoding in Source Reconstructed Brain Space

2024-10-20 · Yonatan Gideoni, Ryan Charles Timms, Oiwi Parker Jones

Non-invasive brainwave decoding is usually done using Magneto/Electroencephalography (MEG/EEG) sensor measurements as inputs. This makes combining datasets and building models with inductive biases difficult as most data…

EEG

Image Ordinal Classification and Understanding: Grid Dropout with Masking Label

2018-05-08 · Chao Zhang, Ce Zhu, Jimin Xiao, Xun Xu 외

Image ordinal classification refers to predicting a discrete target value which carries ordering correlation among image categories. The limited size of labeled ordinal data renders modern deep learning approaches easy t…

Age EstimationClassificationData AugmentationGeneral Classification+1