paper-with-me

Papers

DiffPrompter: Differentiable Implicit Visual Prompts for Semantic-Segmentation in Adverse Conditions

2023-10-06 · Sanket Kalwar, Mihir Ungarala, Shruti Jain, Aaron Monis, Krishna Reddy Konda, Sourav Garg, K Madhava Krishna

Semantic segmentation in adverse weather scenarios is a critical task for autonomous driving systems. While foundation models have shown promise, the need for specialized adaptors becomes evident for handling more challenging scenarios. We introduce DiffPrompter, a novel differentiable visual and latent prompting mechanism aimed at expanding the learning capabilities of existing adaptors in foundation models. Our proposed $\nabla$HFC image processing block excels particularly in adverse weather conditions, where conventional methods often fall short. Furthermore, we investigate the advantages of jointly training visual and latent prompts, demonstrating that this combined approach significantly enhances performance in out-of-distribution scenarios. Our differentiable visual prompts leverage parallel and series architectures to generate prompts, effectively improving object segmentation tasks in adverse conditions. Through a comprehensive series of experiments and evaluations, we provide empirical evidence to support the efficacy of our approach. Project page at https://diffprompter.github.io.

📄 PDF Abstract BibTeX arXiv:2310.04181

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingSegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

Implicit and Explicit Language Guidance for Diffusion-based Visual Perception

2024-04-11 · Hefeng Wang, Jiale Cao, Jin Xie, Aiping Yang 외

Text-to-image diffusion models have shown powerful ability on conditional image synthesis. With large-scale vision-language pre-training, diffusion models are able to generate high-quality images with rich texture and re…

Depth EstimationImage GenerationSemantic Segmentation

EPIPTrack: Rethinking Prompt Modeling with Explicit and Implicit Prompts for Multi-Object Tracking

2025-10-15 · Yukuan Zhang, Jiarui Zhao, Shangqing Nie, Jin Kuang 외 arxiv

Multimodal semantic cues, such as textual descriptions, have shown strong potential in enhancing target perception for tracking. However, existing methods rely on static textual descriptions from large language models, w…

Multi-Object Tracking

Implicit Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis

2023-09-21 · NeurIPS 2023 11

Deep network models are often purely inductive during both training and inference on unseen data. When these models are used for prediction, but they may fail to capture important semantic information and implicit depend…

Cross-Modal RetrievalImage CaptioningImage RetrievalNatural Language Understanding+5

VP3D: Unleashing 2D Visual Prompt for Text-to-3D Generation

2024-03-25 · CVPR 2024 1 · Yang Chen, Yingwei Pan, Haibo Yang, Ting Yao 외

Recent innovations on text-to-3D generation have featured Score Distillation Sampling (SDS), which enables the zero-shot learning of implicit 3D models (NeRF) by directly distilling prior knowledge from 2D diffusion mode…

3D GenerationNeRFText to 3DZero-Shot Learning

Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models

2026-06-24 · Xi Xiao, Xingjian Li, Yunbei Zhang, Cheng Han 외 arxiv

Visual prompt tuning has emerged as a parameter-efficient fine-tuning approach for adapting large-scale Vision Transformers (ViTs) to downstream tasks. As its learnable prompts are applied in input and feature spaces, pr…

parameter-efficient fine-tuningVisual Prompt Tuning