paper-with-me

Papers

Diffusion Probe: Generated Image Result Prediction Using CNN Probes

2026-02-27 · Benlei Cui, Bukun Huang, Zhizeng Ye, Xuemei Dong, Tuo Chen, Hui Xue, Dingkang Yang, Longtao Huang, Jingqun Tang, Haiwen Hong arxiv

Text-to-image (T2I) diffusion models lack an efficient mechanism for early quality assessment, leading to costly trial-and-error in multi-generation scenarios such as prompt iteration, agent-based generation, and flow-grpo. We reveal a strong correlation between early diffusion cross-attention distributions and final image quality. Based on this finding, we introduce Diffusion Probe, a framework that leverages internal cross-attention maps as predictive signals. We design a lightweight predictor that maps statistical properties of early-stage cross-attention extracted from initial denoising steps to the final image's overall quality. This enables accurate forecasting of image quality across diverse evaluation metrics long before full synthesis is complete. We validate Diffusion Probe across a wide range of settings. On multiple T2I models, across early denoising windows, resolutions, and quality metrics, it achieves strong correlation (PCC > 0.7) and high classification performance (AUC-ROC > 0.9). Its reliability translates into practical gains. By enabling early quality-aware decisions in workflows such as prompt optimization, seed selection, and accelerated RL training, the probe supports more targeted sampling and avoids computation on low-potential generations. This reduces computational overhead while improving final output quality.Diffusion Probe is model-agnostic, efficient, and broadly applicable, offering a practical solution for improving T2I generation efficiency through early quality prediction.

📄 PDF Abstract BibTeX arXiv:2602.23783

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pseudorandom Streams within Diffusion Models Act as Learnable Inputs That Affect Generation Quality

2026-08-03 · Shengzhi Deng, Chenqi Ye, Yanze Guo arxiv

Digital learning systems consume concrete pseudorandom values rather than abstract random variables. These values enter the realized loss and its gradient during training. If a pseudorandom stream contains structure that…

Value prediction

TINA+: Probing Residual Visual Knowledge in Unlearned Diffusion Models via Diffusion-Consistent Text-Free Inversion

2026-08-18 · Qianlong Xiang, Miao Zhang, Kun Wang, Haoyu Zhang 외 arxiv

Although text-to-image diffusion models exhibit remarkable generative power, concept erasure techniques are essential for preventing harmful content. Existing adversarial probes evaluate these methods by testing whether …

Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback

2025-07-03 · Nina Konovalova, Maxim Nikolaev, Andrey Kuznetsov, Aibek Alanov arxiv

Despite significant progress in text-to-image diffusion models, achieving precise spatial control over generated outputs remains challenging. ControlNet addresses this by introducing an auxiliary conditioning module, whi…

Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models

2024-11-26 · Colin Conwell, Rupert Tawiah-Quashie, Tomer Ullman

Despite remarkable progress in multi-modal AI research, there is a salient domain in which modern AI continues to lag considerably behind even human children: the reliable deployment of logical operators. Here, we examin…

Text-to-Image Generation

Prediction and Recovery for Adaptive Low-Resolution Person Re-Identification

2020-08-01 · ECCV 2020 8 · Ke Han, Yan Huang, Zerui Chen, Liang Wang 외

Low-resolution person re-identification (LR re-id) is a challenging task with low-resolution probes and high-resolution gallery images. To address the resolution mismatch, existing methods typically recover missing detai…

Person Re-IdentificationSuper-Resolution