paper-with-me

홈 › Papers

Multi-Modal Traffic Sign Detection with Semantic Attributes for Autonomous Driving

2026-08-21 · Meda Lazar, Sourab Sridhar, Shashwata Gupta, Alexandra Tripcea, Varun Ravi, Senthil Yogamani arxiv

Reliable traffic sign detection is a prerequisite for the global deployment of autonomous driving systems, where regulatory compliance and road safety depend on perceiving signs correctly across regions, ranges, and weather conditions. Despite recent progress, vision-based methods continue to face three fundamental limitations: poor cross-regional generalization due to high diversity across countries, degraded performance on small-object detection at long ranges (traffic signs occupy as little as $10{\times}10$ pixels at 200m), and fragile temporal tracking under the strongly non-linear perspective distortion that occurs as a vehicle approaches a sign. In this paper, we address the problem of robust, long-range, region-agnostic traffic sign perception by combining camera and Light Detection and Ranging (LiDAR) sensing. We present a multi-modal detection framework whose Intensity-Aware Deformable Fusion module aligns retro-reflective LiDAR cues with camera features, anchoring detection on geometric invariants rather than region-specific visual appearance. We further introduce a dual motion-model tracker that explicitly accounts for non-linear perspective transformations during vehicle approach, substantially improving temporal consistency over linear motion assumptions. Additionally, we develop a semantic attribute classification pipeline that estimates occlusion level, readability, sign embeddedness, and road relevance, providing actionable context to downstream planning. Extensive evaluation on our dataset, spanning 60+ countries and 2,500+ hours of driving data, shows that the proposed pipeline achieves an Object Miss Ratio (OMR) of 0.49% across 221,068 evaluation sequences, demonstrating globally generalizable traffic sign perception in commercial-grade autonomous driving systems.

📄 PDF Abstract BibTeX arXiv:2608.20874

Code (0)

등록된 구현이 없습니다.

Tasks

Traffic Sign DetectionAutonomous DrivingObject Detection

Similar Papers 제목 키워드 기반

PacketCLIP: Multi-Modal Embedding of Network Traffic and Language for Cybersecurity Reasoning

2025-03-05 · Ryozo Masukawa, Sanggeon Yun, Sungheon Jeong, Wenjun Huang 외

Traffic classification is vital for cybersecurity, yet encrypted traffic poses significant challenges. We present PacketCLIP, a multi-modal framework combining packet data with natural language semantics through contrast…

Anomaly DetectionClassificationGraph Neural NetworkIntrusion Detection+2

FOMO-3D: Using Vision Foundation Models for Long-Tailed 3D Object Detection

2026-03-09 · Anqi Joyce Yang, James Tu, Nikita Dvornik, Enxu Li 외 arxiv

In order to navigate complex traffic environments, self-driving vehicles must recognize many semantic classes pertaining to vulnerable road users or traffic control devices. However, many safety-critical objects (e.g., c…

3D Object Detection

RB-SCD: A New Benchmark for Semantic Change Detection of Roads and Bridges in Traffic Scenes

2025-05-19 · Qingling Shu, Sibao Chen, Zhihui You, Wei Lu 외

Accurate detection of changes in roads and bridges, such as construction, renovation, and demolition, is essential for urban planning and traffic management. However, existing methods often struggle to extract fine-grain…

Change Detection

Contrastive Learning-Driven Traffic Sign Perception: Multi-Modal Fusion of Text and Vision

2025-07-31 · Qiang Lu, Waikit Xiu, Xiying Li, Shenyu Hu 외 arxiv

Traffic sign recognition, as a core component of autonomous driving perception systems, directly influences vehicle environmental awareness and driving safety. Current technologies face two significant challenges: first,…

Traffic Sign RecognitionTraffic Sign DetectionContrastive LearningAutonomous Driving

Multimodal Reasoning with LLM for Encrypted Traffic Interpretation: A Benchmark

2026-04-09 · Longgang Zhang, Xiaowei Fu, Fuxiang Huang, Lei Zhang arxiv

Network traffic, as a key media format, is crucial for ensuring security and communications in modern internet infrastructure. While existing methods offer excellent performance, they face two key bottlenecks: (1) They f…

Multimodal Reasoning