SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation
Personalized image generation aims to faithfully preserve a reference subject's identity while adapting to diverse text prompts. Existing optimization-based methods ensure high fidelity but are computationally expensive, while learning-based approaches offer efficiency at the cost of entangled representations influenced by nuisance factors. We introduce SpotDiff, a novel learning-based method that extracts subject-specific features by spotting and disentangling interference. Leveraging a pre-trained CLIP image encoder and specialized expert networks for pose and background, SpotDiff isolates subject identity through orthogonality constraints in the feature space. To enable principled training, we introduce SpotDiff10k, a curated dataset with consistent pose and background variations. Experiments demonstrate that SpotDiff achieves more robust subject preservation and controllable editing than prior methods, while attaining competitive performance with only 10k training samples.
Code (0)
등록된 구현이 없습니다.
Tasks
Personalized Image GenerationSimilar Papers 제목 키워드 기반
Data Augmentation for Robust Keyword Spotting under Playback Interference
Accurate on-device keyword spotting (KWS) with low false accept and false reject rate is crucial to customer experience for far-field voice control of conversational agents. It is particularly challenging to maintain low…
Acoustic echo cancellationData AugmentationKeyword SpottingYou Only Recognize Once: Towards Fast Video Text Spotting
Video text spotting is still an important research topic due to its various real-applications. Previous approaches usually fall into the four-staged pipeline: text detection in individual images, framewisely recognizing …
Text DetectionText SpottingFacial Expression Spotting Based on Optical Flow Features
The purpose of micro expression (ME) and macro expression (MaE) spotting task is to locate the onset and offset frames of MaE and ME clips. Compared with MaEs, MEs are shorter in duration and lower in intensity, which ma…
Optical Flow EstimationICDAR 2021 Competition on Scene Video Text Spotting
Scene video text spotting (SVTS) is a very important research topic because of many real-life applications. However, only a little effort has put to spotting scene video text, in contrast to massive studies of scene text…
Task 2Text DetectionText SpottingvalidStructural Instability of Feature Composition
Sparse Autoencoders (SAEs) have emerged as a powerful paradigm for disentangling feature superposition in transformer-based architectures, enabling precise control via activation steering. However, the theoretical founda…