paper-with-me

Papers

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

2026-06-26 · Yana Wei, Hongbo Peng, Yanlin Lai, Liang Zhao, Kangheng Lin, En Yu, Keyu Lv, Han Zhou, Yin Tang, Haodong Li, Mitt Huang, Hangyu Guo, Jianjian Sun, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Vishal M. Patel hf

We introduce PerceptionRubrics, a rubric-based evaluation framework that addresses the gap between saturated benchmark scores and real-world brittleness. Shifting evaluation from holistic semantic matching to rigorous atomic auditing, PerceptionRubrics pairs 1,038 information-dense images with over 10,000 instance-specific rubrics. These criteria are derived from golden captions constructed via a novel Circular Peer-Review consensus pipeline and then distilled into a dual-stream system of Must-Right (essential facts) and Easy-Wrong (fine-grained details) rubrics. Crucially, PerceptionRubrics implements a Gated Scoring mechanism: unlike linear averages, failure on mandatory visual facts triggers sharp binary penalties. Extensive evaluation yields critical insights: (1) The Reliability Gap: models often verify fragmented elements correctly yet fail strict conjunctive constraints, exposing brittleness in dense domains; (2) Open-Closed Stratification: contrary to reasoning trends, we reveal a persistent 8% perception deficit between open-source and proprietary frontiers; and (3) Human-Aligned Rigor: our gated metrics substantially out-align conventional benchmarks, validating that strict perceptual fidelity is the prerequisite for reliable generation.

📄 PDF Abstract BibTeX arXiv:2606.28322

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation

2024-06-29 · Jinsheng Huang, Liang Chen, Taian Guo, Fu Zeng 외

Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, often assessed through multiple-choice questions (MCQs) that include an image, a question, and several options. However…

Multiple-choice

Uncertainty-Guided Enhancement on Driving Perception System via Foundation Models

2024-10-02 · Yunhao Yang, Yuxin Hu, Mao Ye, Zaiwei Zhang 외

Multimodal foundation models offer promising advancements for enhancing driving perception systems, but their high computational and financial costs pose challenges. We develop a method that leverages foundation models t…

Conformal PredictionPrediction

BLINK: Multimodal Large Language Models Can See but Not Perceive

2024-04-18 · Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng 외

We introduce Blink, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the Blink tasks can be solved by humans "within a blink" …

Depth EstimationMultiple-choiceVisual Prompting

ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models

2024-10-07 · Ziyue Wang, Chi Chen, Fuwen Luo, Yurui Dong 외

Active perception, a crucial human capability, involves setting a goal based on the current understanding of the environment and performing actions to achieve that goal. Despite significant efforts in evaluating Multimod…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

SAM Audio Judge: A Unified Multimodal Framework for Perceptual Evaluation of Audio Separation

2026-01-27 · Helin Wang, Bowen Shi, Andros Tjandra, John Hoffman 외 arxiv

The performance evaluation remains a complex challenge in audio separation, and existing evaluation metrics are often misaligned with human perception, course-grained, relying on ground truth signals. On the other hand, …