paper-with-me

홈 › Papers

Effective Adaptation in Multi-Task Co-Training for Unified Autonomous Driving

2022-09-19 · Xiwen Liang, Yangxin Wu, Jianhua Han, Hang Xu, Chunjing Xu, Xiaodan Liang

Aiming towards a holistic understanding of multiple downstream tasks simultaneously, there is a need for extracting features with better transferability. Though many latest self-supervised pre-training methods have achieved impressive performance on various vision tasks under the prevailing pretrain-finetune paradigm, their generalization capacity to multi-task learning scenarios is yet to be explored. In this paper, we extensively investigate the transfer performance of various types of self-supervised methods, e.g., MoCo and SimCLR, on three downstream tasks, including semantic segmentation, drivable area segmentation, and traffic object detection, on the large-scale driving dataset BDD100K. We surprisingly find that their performances are sub-optimal or even lag far behind the single-task baseline, which may be due to the distinctions of training objectives and architectural design lied in the pretrain-finetune paradigm. To overcome this dilemma as well as avoid redesigning the resource-intensive pre-training stage, we propose a simple yet effective pretrain-adapt-finetune paradigm for general multi-task training, where the off-the-shelf pretrained models can be effectively adapted without increasing the training overhead. During the adapt stage, we utilize learnable multi-scale adapters to dynamically adjust the pretrained model weights supervised by multi-task objectives while leaving the pretrained knowledge untouched. Furthermore, we regard the vision-language pre-training model CLIP as a strong complement to the pretrain-adapt-finetune paradigm and propose a novel adapter named LV-Adapter, which incorporates language priors in the multi-task model via task-specific prompting and alignment between visual and textual features.

📄 PDF Abstract BibTeX arXiv:2209.08953

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingMulti-Task Learningobject-detectionObject DetectionSemantic SegmentationTraffic Object Detection

Methods 이 논문이 사용한 방법론

Bitcoin Customer Service Number +1-833-534-1729 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Average Pooling 설명 없음
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Random Gaussian Blur Random Gaussian Blur is an image data augmentation technique where we randomly blur the image using a Gaussian distribution. Image Source:…
NT-Xent NT-Xent, or Normalized Temperature-scaled Cross Entropy Loss, is a loss function. Let $\text{sim}\left(\mathbf{u}, \mathbf{v}\right) =…

Similar Papers 제목 키워드 기반

DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles

2025-03-05 · CVPR 2025 1 · Rui Zhao, Weijia Mao, Mike Zheng Shou

Adapting generative models to specific domains presents an effective solution for satisfying specialized requirements. However, adapting to some complex domains remains challenging, especially when these domains require …

Domain AdaptationImage to text

U2A: Unified Unimodal Adaptation for Robust and Efficient Multimodal Learning

2025-01-29 · Md Kaykobad Reza, Niki Nezakati, Ameya Patil, Mashhour Solh 외

Multimodal learning often relies on designing new models and complex training strategies to achieve optimal performance. We present Unified Unimodal Adaptation (U2A), which jointly fine-tunes pretrained unimodal encoders…

UniDA3D: Unified Domain Adaptive 3D Semantic Segmentation Pipeline

2022-12-20 · Ben Fei, Siyuan Huang, Jiakang Yuan, Botian Shi 외

State-of-the-art 3D semantic segmentation models are trained on off-the-shelf public benchmarks, but they will inevitably face the challenge of recognition accuracy drop when these well-trained models are deployed to a n…

3D Semantic SegmentationDomain AdaptationDomain GeneralizationSegmentation+2

TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization

2025-03-25 · CVPR 2025 1 · Liang Pan, Zeshi Yang, Zhiyang Dou, Wenjia Wang 외

Synthesizing diverse and physically plausible Human-Scene Interactions (HSI) is pivotal for both computer animation and embodied AI. Despite encouraging progress, current methods mainly focus on developing separate contr…

ThanoRA: Task Heterogeneity-Aware Multi-Task Low-Rank Adaptation

2025-05-24 · Jian Liang, Wenke Huang, Xianda Guo, Guancheng Wan 외

Low-Rank Adaptation (LoRA) is widely adopted for downstream fine-tuning of foundation models due to its efficiency and zero additional inference cost. Many real-world applications require foundation models to specialize …

Mixture-of-Experts