paper-with-me

홈 › Papers

Seeing Roads Through Words: A Language-Guided Framework for RGB-T Driving Scene Segmentation

2026-02-07 · Ruturaj Reddy, Hrishav Bakul Barua, Junn Yong Loo, Thanh Thi Nguyen, Ganesh Krishnasamy arxiv

Robust semantic segmentation of road scenes under adverse illumination, lighting, and shadow conditions remain a core challenge for autonomous driving applications. RGB-Thermal fusion is a standard approach, yet existing methods apply static fusion strategies uniformly across all conditions, allowing modality-specific noise to propagate throughout the network. Hence, we propose CLARITY that dynamically adapts its fusion strategy to the detected scene condition. Guided by vision-language model (VLM) priors, the network learns to modulate each modality's contribution based on the illumination state while leveraging object embeddings for segmentation, rather than applying a fixed fusion policy. We further introduce two mechanisms - one which preserves valid dark-object semantics that prior noise-suppression methods incorrectly discard, and a hierarchical decoder that enforces structural consistency across scales to sharpen boundaries on thin objects. Experiments on the MFNet dataset demonstrate that CLARITY establishes a new state-of-the-art (SOTA), achieving 62.3% mIoU and 77.5% mAcc.

📄 PDF Abstract BibTeX arXiv:2602.07343

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SegmentationScene SegmentationAutonomous Driving

Similar Papers 제목 키워드 기반

StyleGAN-NADA: CLIP-Guided Domain Adaptation of Image Generators

2021-08-02 · Rinon Gal, Or Patashnik, Haggai Maron, Gal Chechik 외

Can a generative model be trained to produce images from a specific domain, guided by a text prompt only, without seeing any image? In other words: can an image generator be trained "blindly"? Leveraging the semantic pow…

Domain AdaptationImage Manipulation

Seeing the roads through the trees: A benchmark for modeling spatial dependencies with aerial imagery

2024-01-12 · Caleb Robinson, Isaac Corley, Anthony Ortiz, Rahul Dodhia 외

Fully understanding a complex high-resolution satellite or aerial imagery scene often requires spatial reasoning over a broad relevant context. The human object recognition system is able to understand object in a scene …

Object RecognitionRoad SegmentationSemantic SegmentationSpatial Reasoning

Seeing in Words: Learning to Classify through Language Bottlenecks

2023-06-29 · Khalid Saifullah, Yuxin Wen, Jonas Geiping, Micah Goldblum 외

Neural networks for computer vision extract uninterpretable features despite achieving high accuracy on benchmarks. In contrast, humans can explain their predictions using succinct and intuitive descriptions. To incorpor…

Visually grounded cross-lingual keyword spotting in speech

2018-06-13 · Herman Kamper, Michael Roth

Recent work considered how images paired with speech can be used as supervision for building speech systems when transcriptions are not available. We ask whether visual grounding can be used for cross-lingual keyword spo…

Keyword SpottingVisual Grounding

Seeing without Pixels: Perception from Camera Trajectories

2025-11-26 · Zihui Xue, Kristen Grauman, Dima Damen, Andrew Zisserman 외 arxiv

Can one perceive a video's content without seeing its pixels, just from the camera trajectory-the path it carves through space? This paper is the first to systematically investigate this seemingly implausible question. T…

Camera Pose EstimationContrastive Learning