paper-with-me

홈 › Papers

MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation

2026-05-21 · Patryk Bartkowiak, Lennart Petersen, Bartosz Kotrys, Dominik Michels, Soren Pirk, Wojtek Palubicki arxiv

Evaluating single-concept personalization in text-to-image diffusion requires measuring both concept preservation, which captures identity fidelity to a reference, and prompt following, which captures whether the generated scene matches the prompt. Existing metrics commonly compute these signals using global image or text-image embeddings, such as CLIP-I, DINO, and CLIP-T. We show that such metrics correlate poorly with human perception because they attend to the image as a whole instead of separating the concept subject from the background. We introduce MaSC, a masked similarity metric that uses externally provided foreground concept masks to decompose evaluation into subject-specific concept preservation and background-based prompt following. MaSC computes both scores from frozen SigLIP2 SO400M-NaFlex features: concept preservation is measured by masked max-cosine matching between foreground reference patches and generated-image patches, while prompt following is measured by comparing a background-only pooled image embedding to a subject-stripped prompt embedding. On DreamBench++ human ratings, MaSC achieves Krippendorff alpha = 0.471 for concept preservation, outperforming all tested non-LLM baselines and GPT-4V, and approaching GPT-4o. On ORIDa, a real-photo identity-preservation benchmark across physical environments, MaSC achieves AUC = 0.992, nearly perfectly distinguishing same-subject from cross-subject pairs. Its prompt-following score also outperforms the CLIP-T baseline shipped with DreamBench++. These results show that spatially decomposed aggregation is a strong design principle for evaluating concept-driven generation.

📄 PDF Abstract BibTeX arXiv:2605.22469

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SeMaScore : a new evaluation metric for automatic speech recognition tasks

2024-01-15 · Zitha Sasindran, Harsha Yelchuri, T. V. Prabhakar

In this study, we present SeMaScore, generated using a segment-wise mapping and scoring algorithm that serves as an evaluation metric for automatic speech recognition tasks. SeMaScore leverages both the error rate and a …

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Mask to reconstruct: Cooperative Semantics Completion for Video-text Retrieval

2023-05-13 · Han Fang, Zhifei Yang, Xianghao Zang, Chao Ban 외

Recently, masked video modeling has been widely explored and significantly improved the model's understanding ability of visual regions at a local level. However, existing methods usually adopt random masking and follow …

RetrievalText RetrievalVideo RetrievalVideo-Text Retrieval

MASC: Boosting Autoregressive Image Generation with a Manifold-Aligned Semantic Clustering

2025-10-05 · Lixuan He, Shikang Zheng, Linfeng Zhang arxiv

Autoregressive (AR) models have shown great promise in image generation, yet they face a fundamental inefficiency stemming from their core component: a vast, unstructured vocabulary of visual tokens. This conventional ap…

Semantic SimilarityImage Generation

Fast Unlearning at Scale via Margin Self-Correction

2026-06-01 · Federico Di Gennaro, Alexander Shevchenko, Fanny Yang arxiv

Language-model unlearning updates a trained model to behave as if it had not seen selected training examples, while preserving utility and avoiding costly retraining. Existing approaches typically fine-tune the pretraine…

Transcending the "Male Code": Implicit Masculine Biases in NLP Contexts

2023-04-22 · Katie Seaborn, Shruti Chandra, Thibault Fabre

Critical scholarship has elevated the problem of gender bias in data sets used to train virtual assistants (VAs). Most work has focused on explicit biases in language, especially against women, girls, femme-identifying p…

Word Embeddings