paper-with-me

Papers

AnomalyControl: Learning Cross-modal Semantic Features for Controllable Anomaly Synthesis

2024-12-09 · Shidan He, Lei Liu, Xiujun Shu, Bo wang, Yuanhao Feng, Shen Zhao

Anomaly synthesis is a crucial approach to augment abnormal data for advancing anomaly inspection. Based on the knowledge from the large-scale pre-training, existing text-to-image anomaly synthesis methods predominantly focus on textual information or coarse-aligned visual features to guide the entire generation process. However, these methods often lack sufficient descriptors to capture the complicated characteristics of realistic anomalies (e.g., the fine-grained visual pattern of anomalies), limiting the realism and generalization of the generation process. To this end, we propose a novel anomaly synthesis framework called AnomalyControl to learn cross-modal semantic features as guidance signals, which could encode the generalized anomaly cues from text-image reference prompts and improve the realism of synthesized abnormal samples. Specifically, AnomalyControl adopts a flexible and non-matching prompt pair (i.e., a text-image reference prompt and a targeted text prompt), where a Cross-modal Semantic Modeling (CSM) module is designed to extract cross-modal semantic features from the textual and visual descriptors. Then, an Anomaly-Semantic Enhanced Attention (ASEA) mechanism is formulated to allow CSM to focus on the specific visual patterns of the anomaly, thus enhancing the realism and contextual relevance of the generated anomaly features. Treating cross-modal semantic features as the prior, a Semantic Guided Adapter (SGA) is designed to encode effective guidance signals for the adequate and controllable synthesis process. Extensive experiments indicate that AnomalyControl can achieve state-of-the-art results in anomaly synthesis compared with existing methods while exhibiting superior performance for downstream tasks.

📄 PDF Abstract BibTeX arXiv:2412.06510

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Focus 설명 없음
Adapter 설명 없음

Similar Papers 제목 키워드 기반

RoomPilot: Controllable Indoor Scene Synthesis via Multimodal Semantic Parsing

2025-12-12 · Wentang Chen, Shougao Zhang, Yiman Zhang, Tianhao Zhou 외 arxiv

Generating controllable indoor scenes is fundamental to applications in game development, architectural visualization, and embodied AI. However, existing approaches either support a limited input modalities or rely on im…

Indoor Scene SynthesisScene GenerationSemantic Parsing

PMMD: A pose-guided multi-view multi-modal diffusion for person generation

2025-12-17 · Ziyu Shang, Haoran Liu, Rongchao Zhang, Zhiqian Wei 외 arxiv

Generating consistent human images with controllable pose and appearance is essential for applications in virtual try on, image editing, and digital human creation. Current methods often suffer from occlusions, garment s…

Image Editing

EmoLat: Text-driven Image Sentiment Transfer via Emotion Latent Space

2026-01-17 · Jing Zhang, Bingjie Fan, Jixiang Zhu, Zhe Wang arxiv

We propose EmoLat, a novel emotion latent space that enables fine-grained, text-driven image sentiment transfer by modeling cross-modal correlations between textual semantics and visual emotion features. Within EmoLat, a…

Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering

2026-07-29 · Yu Wang, Sharon Li arxiv

Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet it remains unclear whether these capabilities share a unified and transferable semantic space. This questio…

MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face Generation

2026-03-30 · Bharath Krishnamurthy, Ajita Rattani arxiv

Recent multimodal face generation models address the spatial control limitations of text-to-image diffusion models by augmenting text-based conditioning with spatial priors such as segmentation masks, sketches, or edge m…