paper-with-me

홈 › Papers

Envisage: Diffusion-Based Rhinoplasty Goal Visualization with Mask-Decomposed Evaluation

2026-06-26 · Mudit Agarwal, Amit D. Bhrany arxiv

Localized generative editing needs localized evaluation: full-image identity metrics are structurally confounded under hard-composited edits. We present Envisage, a FLUX.1-Fill inpainting reference pipeline for rhinoplasty goal visualization from a single frontal photograph. The pipeline combines 8 rhinoplasty clinical presets (the released framework also includes 8 blepharoplasty and 8 rhytidectomy presets), MediaPipe masks, and hard-mask compositing. The composite preserves outside-mask pixels by construction, so full-face identity scores are dominated by copied pixels rather than by the diffusion backbone. Because full-face identity metrics cannot grade localized edits, we introduce SurgicalScore, a mask-decomposed 0-1 protocol scoring edit direction, edit magnitude, masked LPIPS, realism, and outside-mask preservation; SS_raw assigns 0.919 [0.918, 0.920] to a perfect-predictor control , anchoring the ceiling. On N=211, the paired ArcFace gain (output-to-GT minus input-to-GT) is negative for all methods (Envisage -0.048 smallest, vs. ICEdit -0.139, Kontext -0.242, InstructPix2Pix -0.294; p < 1e-4), with external validation on a 457-pair ASPS/PCA corpus showing a larger negative gap. With SurgicalScore, Envisage achieves the highest score (0.599 [0.579, 0.619]) and leads on both metrics, but the all-negative ArcFace gap shows that full-face identity is poorly aligned with localized surgical accuracy under hard compositing. A 5-seed GT-oracle (an upper bound, not a deployable result) reduces the residual ArcFace gap by 73% (-0.054 to -0.015), with positive output-to-GT gain on 33.9% of cases, indicating candidate-space headroom for a learned ranker. For localized edits, progress should be measured with edit-region fidelity rather than full-face identity metrics. We release Envisage, SurgicalScore, preset definitions, and matched split manifests.

📄 PDF Abstract BibTeX arXiv:2606.28628

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion

2024-07-17 · Huiguo He, Huan Yang, Zixi Tuo, Yuan Zhou 외

Story visualization aims to create visually compelling images or videos corresponding to textual narratives. Despite recent advances in diffusion models yielding promising results, existing methods still struggle to crea…

DescriptiveStory Visualization

Automated Assessment of Aesthetic Outcomes in Facial Plastic Surgery

2025-08-18 · Pegah Varghaei, Kiran Abraham-Aggarwal, Manoj T. Abraham, Arun Ross arxiv

We introduce a scalable, interpretable computer-vision framework for quantifying aesthetic outcomes of facial plastic surgery using frontal photographs. Our pipeline leverages automated landmark detection, geometric faci…

Age Estimation

Self-Supervised Learning of Time Series Representation via Diffusion Process and Imputation-Interpolation-Forecasting Mask

2024-05-09 · Zineb Senane, Lele Cao, Valentin Leonhard Buchner, Yusuke Tashiro 외

Time Series Representation Learning (TSRL) focuses on generating informative representations for various Time Series (TS) modeling tasks. Traditional Self-Supervised Learning (SSL) methods in TSRL fall into four main cat…

Anomaly DetectionImputationRepresentation LearningSelf-Supervised Learning+1

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models

2026-06-01 · Yingzi Ma, Zhengyue Zhao, Xiaogeng Liu, Minhui Xue 외 arxiv

Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autoregressive LLMs. Because mask tokens are …

CinePreGen: Camera Controllable Video Previsualization via Engine-powered Diffusion

2024-08-30 · Yiran Chen, Anyi Rao, Xuekun Jiang, Shishi Xiao 외

With advancements in video generative AI models (e.g., SORA), creators are increasingly using these techniques to enhance video previsualization. However, they face challenges with incomplete and mismatched AI workflows.…