paper-with-me

홈 › Papers

Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)

2024-04-05 · Michael Saxon, Fatima Jahara, Mahsa Khoshnoodi, Yujie Lu, Aditya Sharma, William Yang Wang

With advances in the quality of text-to-image (T2I) models has come interest in benchmarking their prompt faithfulness -- the semantic coherence of generated images to the prompts they were conditioned on. A variety of T2I faithfulness metrics have been proposed, leveraging advances in cross-modal embeddings and vision-language models (VLMs). However, these metrics are not rigorously compared and benchmarked, instead presented with correlation to human Likert scores over a set of easy-to-discriminate images against seemingly weak baselines. We introduce T2IScoreScore, a curated set of semantic error graphs containing a prompt and a set of increasingly erroneous images. These allow us to rigorously judge whether a given prompt faithfulness metric can correctly order images with respect to their objective error count and significantly discriminate between different error nodes, using meta-metric scores derived from established statistical tests. Surprisingly, we find that the state-of-the-art VLM-based metrics (e.g., TIFA, DSG, LLMScore, VIEScore) we tested fail to significantly outperform simple (and supposedly worse) feature-based metrics like CLIPScore, particularly on a hard subset of naturally-occurring T2I model errors. TS2 will enable the development of better T2I prompt faithfulness metrics through more rigorous comparison of their conformity to expected orderings and separations under objective criteria.

📄 PDF Abstract BibTeX arXiv:2404.04251

Code (1)

michaelsaxon/T2IScoreScore 공식 구현 pytorch

Tasks

Benchmarking

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension

2024-09-23 · Junzhuo Liu, Xuzheng Yang, Weiwei Li, Peng Wang

Referring Expression Comprehension (REC) is a crucial cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and language-to-image grounding. Consequently, it serves …

Image ComprehensionReferring ExpressionReferring Expression ComprehensionVisual Reasoning

RAVE: Retrieval and Scoring Aware Verifiable Claim Detection

2025-09-19 · Yufeng Li, Arkaitz Zubiaga arxiv

The rapid spread of misinformation on social media underscores the need for scalable fact-checking tools. A key step is claim detection, which identifies statements that can be objectively verified. Prior approaches ofte…

RoBIC: A benchmark suite for assessing classifiers robustness

2021-02-10 · Thibault Maho, Benoît Bonnet, Teddy Furon, Erwan Le Merrer

Many defenses have emerged with the development of adversarial attacks. Models must be objectively evaluated accordingly. This paper systematically tackles this concern by proposing a new parameter-free benchmark we coin…

Adversarial Attack

CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization

2025-08-12 · Xinge Ye, Rui Wang, Yuchuan Wu, Victor Ma 외 arxiv

Reinforcement Learning Fine-Tuning (RLFT) has achieved notable success in tasks with objectively verifiable answers (e.g., code generation, mathematical reasoning), yet struggles with open-ended subjective tasks like rol…

Reinforcement LearningMathematical ReasoningCode Generation

VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement Learning

2025-08-08 · Linhan Cao, Wei Sun, Weixia Zhang, Xiangyang Zhu 외 arxiv

Video quality assessment (VQA) aims to objectively quantify perceptual quality degradation in alignment with human visual perception. Despite recent advances, existing VQA models still suffer from two critical limitation…

Video Quality AssessmentReinforcement Learning