paper-with-me

홈 › Papers

VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation

2025-08-08 · Kaiyuan Jiang, Ruoxi Sun, Ying Cao, Yuqi Xu, Xinran Zhang, Junyan Guo, ChengSheng Deng arxiv

We present VISTAR, a user-centric, multi-dimensional benchmark for text-to-image (T2I) evaluation that addresses the limitations of existing metrics. VISTAR introduces a two-tier hybrid paradigm: it employs deterministic, scriptable metrics for physically quantifiable attributes (e.g., text rendering, lighting) and a novel Hierarchical Weighted P/N Questioning (HWPQ) scheme that uses constrained vision-language models to assess abstract semantics (e.g., style fusion, cultural fidelity). Grounded in a Delphi study with 120 experts, we defined seven user roles and nine evaluation angles to construct the benchmark, which comprises 2,845 prompts validated by over 15,000 human pairwise comparisons. Our metrics achieve high human alignment (>75%), with the HWPQ scheme reaching 85.9% accuracy on abstract semantics, significantly outperforming VQA baselines. Comprehensive evaluation of state-of-the-art models reveals no universal champion, as role-weighted scores reorder rankings and provide actionable guidance for domain-specific deployment. All resources are publicly released to foster reproducible T2I assessment.

📄 PDF Abstract BibTeX arXiv:2508.06152

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Visually Interpretable Subtask Reasoning for Visual Question Answering

2025-05-12 · Yu Cheng, Arushi Goel, Hakan Bilen

Answering complex visual questions like `Which red furniture can be used for sitting?' requires multi-step reasoning, including object recognition, attribute filtering, and relational understanding. Recent work improves …

AttributeObject RecognitionQuestion AnsweringVisual Question Answering

VistaRef: Boosting Visual Spatial Orientation Awareness for Pointing-to-Object Detection

2026-06-23 · Ling Li, Zhizhen Cai, Xinkun Wu, Ziyu Zhu 외 arxiv

Grounding deictic gestures in natural images is fundamental to AR and human-robot collaboration, providing a basis for seamless spatial interaction. While Transformer-based visual models have achieved significant progres…

Object Detection

RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing

2025-07-27 · Hao Xiang, Tianyi Tang, Yang Su, Bowen Yu 외 arxiv

Recent advancements in Large Language Models (LLMs) have shown outstanding potential for role-playing applications. Evaluating these capabilities is becoming crucial yet remains challenging. Existing benchmarks mostly ad…

Transparent Adaptive Learning via Data-Centric Multimodal Explainable AI

2025-08-01 · Maryam Mosleh, Marie Devlin, Ellis Solaiman arxiv

Artificial intelligence-driven adaptive learning systems are reshaping education through data-driven adaptation of learning experiences. Yet many of these systems lack transparency, offering limited insight into how deci…

EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning

2026-05-17 · Zeyu Wang, Chang Liu, Eduardus Tjitrahardja, Yuntao Wang 외 arxiv

Despite extensive efforts on egocentric video datasets and benchmarks, understanding users' internal states, which is crucial for enabling seamless AI assistant experiences, remains largely overlooked. In this work, we i…