paper-with-me

Papers

CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks

2025-05-16 · Christoph Leiter, Yuki M. Asano, Margret Keuper, Steffen Eger

The assessment of evaluation metrics (meta-evaluation) is crucial for determining the suitability of existing metrics in text-to-image (T2I) generation tasks. Human-based meta-evaluation is costly and time-intensive, and automated alternatives are scarce. We address this gap and propose CROC: a scalable framework for automated Contrastive Robustness Checks that systematically probes and quantifies metric robustness by synthesizing contrastive test cases across a comprehensive taxonomy of image properties. With CROC, we generate a pseudo-labeled dataset (CROC$^{syn}$) of over one million contrastive prompt-image pairs to enable a fine-grained comparison of evaluation metrics. We also use the dataset to train CROCScore, a new metric that achieves state-of-the-art performance among open-source methods, demonstrating an additional key application of our framework. To complement this dataset, we introduce a human-supervised benchmark (CROC$^{hum}$) targeting especially challenging categories. Our results highlight robustness issues in existing metrics: for example, many fail on prompts involving negation, and all tested open-source metrics fail on at least 25% of cases involving correct identification of body parts.

📄 PDF Abstract BibTeX arXiv:2505.11314

Code (1)

gringham/croc 공식 구현 pytorch

Tasks

Negation

Similar Papers 제목 키워드 기반

A Macrocolumn Architecture Implemented with Spiking Neurons

2022-07-11 · James E. Smith

The macrocolumn is a key component of a neuromorphic computing system that interacts with an external environment under control of an agent. Environments are learned and stored in the macrocolumn as labeled directed grap…

Navigate

Fourier neural operator for real-time simulation of 3D dynamic urban microclimate

2023-08-08 · Wenhui Peng, Shaoxiang Qin, Senwen Yang, Jianchun Wang 외

Global urbanization has underscored the significance of urban microclimates for human comfort, health, and building/urban energy efficiency. They profoundly influence building design and urban planning as major environme…

Beyond Simulation: Benchmarking World Models for Planning and Causality in Autonomous Driving

2025-08-03 · Hunter Schofield, Mohammed Elmahgiubi, Kasra Rezaee, Jinjun Shan arxiv

World models have become increasingly popular in acting as learned traffic simulators. Recent work has explored replacing traditional traffic simulators with world models for policy training. In this work, we explore the…

Autonomous Driving

CroCoSum: A Benchmark Dataset for Cross-Lingual Code-Switched Summarization

2023-03-07 · Ruochen Zhang, Carsten Eickhoff

Cross-lingual summarization (CLS) has attracted increasing interest in recent years due to the availability of large-scale web-mined datasets and the advancements of multilingual language models. However, given the raren…

Articles

microCLIP: Unsupervised CLIP Adaptation via Coarse-Fine Token Fusion for Fine-Grained Image Classification

2025-10-02 · Sathira Silva, Eman Ali, Chetan Arora, Muhammad Haris Khan arxiv

Unsupervised adaptation of CLIP-based vision-language models (VLMs) for fine-grained image classification requires sensitivity to microscopic local cues. While CLIP exhibits strong zero-shot transfer, its reliance on coa…

Fine-Grained Image Classification