paper-with-me

홈 › Papers

Fast Scene Understanding for Autonomous Driving

2017-08-08 · Davy Neven, Bert de Brabandere, Stamatios Georgoulis, Marc Proesmans, Luc van Gool

Most approaches for instance-aware semantic labeling traditionally focus on accuracy. Other aspects like runtime and memory footprint are arguably as important for real-time applications such as autonomous driving. Motivated by this observation and inspired by recent works that tackle multiple tasks with a single integrated architecture, in this paper we present a real-time efficient implementation based on ENet that solves three autonomous driving related tasks at once: semantic scene segmentation, instance segmentation and monocular depth estimation. Our approach builds upon a branched ENet architecture with a shared encoder but different decoder branches for each of the three tasks. The presented method can run at 21 fps at a resolution of 1024x512 on the Cityscapes dataset without sacrificing accuracy compared to running each task separately.

📄 PDF Abstract BibTeX arXiv:1708.02550

Code (1)

davyneven/fastSceneUnderstanding tf

Tasks

Autonomous DrivingDecoderDepth EstimationInstance SegmentationMonocular Depth EstimationScene SegmentationScene UnderstandingSegmentationSemantic Segmentation

Methods 이 논문이 사용한 방법론

Dilated Convolution 설명 없음
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Batch Normalization 설명 없음
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
ENet Dilated Bottleneck ENet Dilated Bottleneck is an image model block used in the ENet semantic segmentation architecture. It is the same as a regular…
ENet Bottleneck ENet Bottleneck is an image model block used in the ENet semantic segmentation architecture. Each block consists of three…
ENet Initial Block The ENet Initial Block is an image model block used in the ENet semantic segmentation architecture. [Max…

Similar Papers 제목 키워드 기반

VAD: Vectorized Scene Representation for Efficient Autonomous Driving

2023-03-21 · ICCV 2023 1 · Bo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao 외

Autonomous driving requires a comprehensive understanding of the surrounding environment for reliable trajectory planning. Previous works rely on dense rasterized scene representation (e.g., agent occupancy and semantic …

Autonomous DrivingBench2DriveTrajectory Planning

Predicting the Road Ahead: A Knowledge Graph based Foundation Model for Scene Understanding in Autonomous Driving

2025-03-24 · Hongkuan Zhou, Stefan Schmid, Yicong Li, Lavdim Halilaj 외

The autonomous driving field has seen remarkable advancements in various topics, such as object recognition, trajectory prediction, and motion planning. However, current approaches face limitations in effectively compreh…

Autonomous DrivingKnowledge GraphsMotion PlanningObject Recognition+2

GaussianPretrain: A Simple Unified 3D Gaussian Representation for Visual Pre-training in Autonomous Driving

2024-11-19 · Shaoqing Xu, Fang Li, Shengyin Jiang, Ziying Song 외

Self-supervised learning has made substantial strides in image processing, while visual pre-training for autonomous driving is still in its infancy. Existing methods often focus on learning geometric scene information wh…

3D Object DetectionAutonomous DrivingGPUNeRF+4

LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving

2025-08-17 · Nan Song, Bozhou Zhang, Xiatian Zhu, Jiankang Deng 외 arxiv

Large vision-language models (VLMs) have shown promising capabilities in scene understanding, enhancing the explainability of driving behaviors and interactivity with users. Existing methods primarily fine-tune VLMs on o…

Scene UnderstandingAutonomous DrivingScene Recognition

VLP: Vision Language Planning for Autonomous Driving

2024-01-10 · CVPR 2024 1 · Chenbin Pan, Burhaneddin Yaman, Tommaso Nesti, Abhirup Mallik 외

Autonomous driving is a complex and challenging task that aims at safe motion planning through scene understanding and reasoning. While vision-only autonomous driving methods have recently achieved notable performance, t…

Autonomous DrivingMotion PlanningScene Understanding