paper-with-me

홈 › Papers

Multi-stream CNN based Video Semantic Segmentation for Automated Driving

2019-01-08 · Ganesh Sistu, Sumanth Chennupati, Senthil Yogamani

Majority of semantic segmentation algorithms operate on a single frame even in the case of videos. In this work, the goal is to exploit temporal information within the algorithm model for leveraging motion cues and temporal consistency. We propose two simple high-level architectures based on Recurrent FCN (RFCN) and Multi-Stream FCN (MSFCN) networks. In case of RFCN, a recurrent network namely LSTM is inserted between the encoder and decoder. MSFCN combines the encoders of different frames into a fused encoder via 1x1 channel-wise convolution. We use a ResNet50 network as the baseline encoder and construct three networks namely MSFCN of order 2 & 3 and RFCN of order 2. MSFCN-3 produces the best results with an accuracy improvement of 9% and 15% for Highway and New York-like city scenarios in the SYNTHIA-CVPR'16 dataset using mean IoU metric. MSFCN-3 also produced 11% and 6% for SegTrack V2 and DAVIS datasets over the baseline FCN network. We also designed an efficient version of MSFCN-2 and RFCN-2 using weight sharing among the two encoders. The efficient MSFCN-2 provided an improvement of 11% and 5% for KITTI and SYNTHIA with negligible increase in computational complexity compared to the baseline version.

📄 PDF Abstract BibTeX arXiv:1901.02511

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSemantic SegmentationVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
FCN Fully Convolutional Networks, or FCNs, are an architecture used mainly for semantic segmentation. They employ solely locally connected layers, such as…
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

DALES: Automated Tool for Detection, Annotation, Labelling, and Segmentation of Multiple Objects in Multi-Camera Video Streams

2014-08-01 · WS 2014 8 · Mohammad Bhat, Joanna Isabelle Olszewska

Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features

2025-10-19 · Shihao Ji, Zihui Song arxiv

The remarkable zero-shot reasoning capabilities of large-scale Visual Language Models (VLMs) on static images have yet to be fully translated to the video domain. Conventional video understanding models often rely on ext…

Unsupervised Semantic Scene Labeling for Streaming Data

2017-07-01 · CVPR 2017 7 · Maggie Wigness, John G. Rogers III

We introduce an unsupervised semantic scene labeling approach that continuously learns and adapts semantic models discovered within a data stream. While closely related to unsupervised video segmentation, our algorithm i…

Scene LabelingSegmentationVideo SegmentationVideo Semantic Segmentation

SAM4D: Segment Anything in Camera and LiDAR Streams

2025-06-26 · Jianyun Xu, Song Wang, Ziqian Ni, Chunyong Hu 외

We present SAM4D, a multi-modal and temporal foundation model designed for promptable segmentation across camera and LiDAR streams. Unified Multi-modal Positional Encoding (UMPE) is introduced to align camera and LiDAR f…

4D reconstructionAutonomous DrivingMotion CompensationSegmentation

VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs

2025-12-23 · Naishan Zheng, Jie Huang, Qingpei Guo, Feng Zhao arxiv

Understanding long videos with multimodal large language models (MLLMs) remains challenging due to the heavy redundancy across frames and the need for temporally coherent representations. Existing static strategies, such…

Event Segmentation