paper-with-me

Papers

ED-SAM: An Efficient Diffusion Sampling Approach to Domain Generalization in Vision-Language Foundation Models

2024-06-03 · Thanh-Dat Truong, Xin Li, Bhiksha Raj, Jackson Cothren, Khoa Luu

The Vision-Language Foundation Model has recently shown outstanding performance in various perception learning tasks. The outstanding performance of the vision-language model mainly relies on large-scale pre-training datasets and different data augmentation techniques. However, the domain generalization problem of the vision-language foundation model needs to be addressed. This problem has limited the generalizability of the vision-language foundation model to unknown data distributions. In this paper, we introduce a new simple but efficient Diffusion Sampling approach to Domain Generalization (ED-SAM) to improve the generalizability of the vision-language foundation model. Our theoretical analysis in this work reveals the critical role and relation of the diffusion model to domain generalization in the vision-language foundation model. Then, based on the insightful analysis, we introduce a new simple yet effective Transport Transformation to diffusion sampling method. It can effectively generate adversarial samples to improve the generalizability of the foundation model against unknown data distributions. The experimental results on different scales of vision-language pre-training datasets, including CC3M, CC12M, and LAION400M, have consistently shown State-of-the-Art performance and scalability of the proposed ED-SAM approach compared to the other recent methods.

📄 PDF Abstract BibTeX arXiv:2406.01432

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationDomain GeneralizationLanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

SDPose: Exploiting Diffusion Priors for Out-of-Domain and Robust Pose Estimation

2025-09-29 · Shuang Liang, Jing He, Chuanmeizhi Wang, Lejun Liao 외 arxiv

Pre-trained diffusion models provide rich latent features across U-Net levels and are emerging as powerful vision backbones. While prior works such as Marigold and Lotus repurpose diffusion priors for dense geometric per…

Domain GeneralizationPose Estimation

Frequency-Aware Vision-Language Multimodality Generalization Network for Remote Sensing Image Classification

2025-11-13 · Junjie Zhang, Feng Zhao, Hanqiang Liu, Jun Yu arxiv

The booming remote sensing (RS) technology is giving rise to a novel multimodality generalization task, which requires the model to overcome data heterogeneity while possessing powerful cross-scene generalization ability…

Remote Sensing Image Classification

Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference

2026-07-14 · Yuzhou Wu, Yuxin Zheng, Muchun Niu, Yishan Yang 외 arxiv

Vision-Language-Action (VLA) models exhibit strong generalization for robotic manipulation, yet their high inference latency limits real time deployment. We identify two primary sources of temporal redundancy in existing…

MFVLR: Multi-domain Fine-grained Vision-Language Reconstruction for Generalizable Diffusion Face Forgery Detection and Localization

2026-05-11 · Yaning Zhang, Tianyi Wang, Zan Gao, Yibo Zhao 외 arxiv

The swift advancement in photo-realistic face generation technology has sparked considerable concerns across society and academia, emphasizing the requirement of generalizable face forgery detection and localization meth…

Representation Learning

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

2025-07-07 · Juyi Lin, Amir Taherin, Arash Akbari, Arman Akbari 외

Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language. However, their generalization remains limited when applied to novel objects…

Depth EstimationVision-Language-Action