paper-with-me

홈 › Papers

Measuring What Matters: Synthetic Benchmarks for Concept Bottleneck Models

2026-06-03 · Julian Skirzynski, Harry Cheon, Shreyas Kadekodi, Meredith Stewart, Berk Ustun arxiv

Concept bottleneck models predict outcomes from high-level concepts detected in inputs. Although concepts provide a simple way to reap benefits from interpretability, very few datasets include concept labels. This limits researchers' ability to determine which problems are suitable for these models, isolate the factors that drive their performance or lead to failures, or uncover which algorithms perform well. In this paper, we develop synthetic benchmarks for concept-bottleneck models, focusing on their two main use cases: decision support, in which models assist humans in making better decisions, and automation, in which models handle routine tasks without supervision. Our benchmarks can generate labeled datasets while controlling for properties that affect performance, including data modality, concept choice, annotation quality, and completeness. We demonstrate how the benchmarks can be used to evaluate representative classes of concept bottleneck models. Our demonstrations show how the benchmarks can diagnose failure modes and guide follow-up testing.

📄 PDF Abstract BibTeX arXiv:2606.04326

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

2025-11-03 · Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner 외 arxiv

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as 'safety'…

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare

2026-05-08 · Prasanna Desikan, Harshit Rajgarhia, Shivali Dalmia, Ananya Mantravadi arxiv

AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard training and validation datasets were never designed to capture. Eval…

Gaussian Process Probes (GPP) for Uncertainty-Aware Probing

2023-09-21 · NeurIPS 2023 11

Understanding which concepts models can and cannot represent has been fundamental to many tasks: from effective and responsible use of models to detecting out of distribution data. We introduce Gaussian process probes (G…

Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks

2026-07-06 · Yibo Hu, Jiaming Qu arxiv

LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response. We show that most of this apparent conformity survives even after the peer is removed. The reason is …

What Makes a Concept Complex? Measuring Conceptual Complexity as a Precursor for Text Simplification

2021-07-01 · TRITON 2021 7 · Anne Eschenbruecher

Advancements within the field of text simplification (TS) have primarily been within syntactic or lexical simplification. However, conceptual simplification has previously been identified as another field of TS that has …

Binary ClassificationLexical SimplificationReading ComprehensionText Simplification