paper-with-me

홈 › Papers

ConceptMix++: Leveling the Playing Field in Text-to-Image Benchmarking via Iterative Prompt Optimization

2025-07-04 · Haosheng Gan, Berk Tinaz, Mohammad Shahab Sepehri, Zalan Fabian, Mahdi Soltanolkotabi arxiv

Current text-to-image (T2I) benchmarks evaluate models on rigid prompts, potentially underestimating true generative capabilities due to prompt sensitivity and creating biases that favor certain models while disadvantaging others. We introduce ConceptMix++, a framework that disentangles prompt phrasing from visual generation capabilities by applying iterative prompt optimization. Building on ConceptMix, our approach incorporates a multimodal optimization pipeline that leverages vision-language model feedback to refine prompts systematically. Through extensive experiments across multiple diffusion models, we show that optimized prompts significantly improve compositional generation performance, revealing previously hidden model capabilities and enabling fairer comparisons across T2I models. Our analysis reveals that certain visual concepts -- such as spatial relationships and shapes -- benefit more from optimization than others, suggesting that existing benchmarks systematically underestimate model performance in these categories. Additionally, we find strong cross-model transferability of optimized prompts, indicating shared preferences for effective prompt phrasing across models. These findings demonstrate that rigid benchmarking approaches may significantly underrepresent true model capabilities, while our framework provides more accurate assessment and insights for future development.

📄 PDF Abstract BibTeX arXiv:2507.03275

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty

2024-08-26 · Xindi Wu, Dingli Yu, Yangsibo Huang, Olga Russakovsky 외

Compositionality is a critical capability in Text-to-Image (T2I) models, as it reflects their ability to understand and combine multiple concepts from text descriptions. Existing evaluations of compositional capability r…

DiversityImage Generation

Human-AI Collaborative Bot Detection in MMORPGs

2025-08-28 · Jaeman Son, Hyunsoo Kim arxiv

In Massively Multiplayer Online Role-Playing Games (MMORPGs), auto-leveling bots exploit automated programs to level up characters at scale, undermining gameplay balance and fairness. Detecting such bots is challenging, …

Representation Learning

Leveling the Playing Field -- Fairness in AI Versus Human Game Benchmarks

2019-03-17 · Rodrigo Canaan, Christoph Salge, Julian Togelius, Andy Nealen

From the beginning if the history of AI, there has been interest in games as a platform of research. As the field developed, human-level competence in complex games became a target researchers worked to reach. Only relat…

Fairness

Is Deep Reinforcement Learning Really Superhuman on Atari? Leveling the playing field

2019-08-13 · Marin Toromanoff, Emilie Wirbel, Fabien Moutarde

Consistent and reproducible evaluation of Deep Reinforcement Learning (DRL) is not straightforward. In the Arcade Learning Environment (ALE), small changes in environment parameters such as stochasticity or the maximum a…

Atari GamesDeep Reinforcement LearningGeneral Reinforcement Learningreinforcement-learning+2

Leveling3D: Leveling Up 3D Reconstruction with Feed-Forward 3D Gaussian Splatting and Geometry-Aware Generation

2026-03-17 · Yiming Huang, Baixiang Huang, Beilei Cui, Chi Kit Ng 외 arxiv

Feed-forward 3D reconstruction has revolutionized 3D vision, providing a powerful baseline for downstream tasks such as novel-view synthesis with 3D Gaussian Splatting. Previous works explore fixing the corrupted renderi…

3D ReconstructionDepth Estimation