paper-with-me

Papers

Prompt-Driven Feature Diffusion for Open-World Semi-Supervised Learning

2024-04-17 · Marzi Heidari, Hanping Zhang, Yuhong Guo

In this paper, we present a novel approach termed Prompt-Driven Feature Diffusion (PDFD) within a semi-supervised learning framework for Open World Semi-Supervised Learning (OW-SSL). At its core, PDFD deploys an efficient feature-level diffusion model with the guidance of class-specific prompts to support discriminative feature representation learning and feature generation, tackling the challenge of the non-availability of labeled data for unseen classes in OW-SSL. In particular, PDFD utilizes class prototypes as prompts in the diffusion model, leveraging their class-discriminative and semantic generalization ability to condition and guide the diffusion process across all the seen and unseen classes. Furthermore, PDFD incorporates a class-conditional adversarial loss for diffusion model training, ensuring that the features generated via the diffusion process can be discriminatively aligned with the class-conditional features of the real data. Additionally, the class prototypes of the unseen classes are computed using only unlabeled instances with confident predictions within a semi-supervised learning framework. We conduct extensive experiments to evaluate the proposed PDFD. The empirical results show PDFD exhibits remarkable performance enhancements over many state-of-the-art existing methods.

📄 PDF Abstract BibTeX arXiv:2404.11795

Code (0)

등록된 구현이 없습니다.

Tasks

Open-World Semi-Supervised LearningRepresentation Learning

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

DiffusionGPT: LLM-Driven Text-to-Image Generation System

2024-01-18 · Jie Qin, Jie Wu, Weifeng Chen, Yuxi Ren 외

Diffusion models have opened up new avenues for the field of image generation, resulting in the proliferation of high-quality models shared on open-source platforms. However, a major challenge persists in current text-to…

Image GenerationModel SelectionText to Image GenerationText-to-Image Generation

MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation

2024-04-08 · Kunpeng Song, Yizhe Zhu, Bingchen Liu, Qing Yan 외

In this paper, we present MoMA: an open-vocabulary, training-free personalized image model that boasts flexible zero-shot capabilities. As foundational text-to-image models rapidly evolve, the demand for robust image-to-…

Image GenerationImage-to-Image TranslationLanguage ModelingLanguage Modelling+3

Towards Training-free Open-world Segmentation via Image Prompt Foundation Models

2023-10-17 · Lv Tang, Peng-Tao Jiang, Hao-Ke Xiao, Bo Li

The realm of computer vision has witnessed a paradigm shift with the advent of foundational models, mirroring the transformative influence of large language models in the domain of natural language processing. This paper…

Segmentation

OpenSDI: Spotting Diffusion-Generated Images in the Open World

2025-03-25 · CVPR 2025 1 · Yabin Wang, Zhiwu Huang, Xiaopeng Hong

This paper identifies OpenSDI, a challenge for spotting diffusion-generated images in open-world settings. In response to this challenge, we define a new benchmark, the OpenSDI dataset (OpenSDID), which stands out from e…

Lightweight Language-driven Grasp Detection using Conditional Consistency Model

2024-07-25 · Nghia Nguyen, Minh Nhat Vu, Baoru Huang, An Vuong 외

Language-driven grasp detection is a fundamental yet challenging task in robotics with various industrial applications. In this work, we present a new approach for language-driven grasp detection that leverages the conce…

Denoising