paper-with-me

홈 › Papers

Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model Behavior

2021-07-20 · Angie Boggust, Benjamin Hoover, Arvind Satyanarayan, Hendrik Strobelt

Saliency methods -- techniques to identify the importance of input features on a model's output -- are a common step in understanding neural network behavior. However, interpreting saliency requires tedious manual inspection to identify and aggregate patterns in model behavior, resulting in ad hoc or cherry-picked analysis. To address these concerns, we present Shared Interest: metrics for comparing model reasoning (via saliency) to human reasoning (via ground truth annotations). By providing quantitative descriptors, Shared Interest enables ranking, sorting, and aggregating inputs, thereby facilitating large-scale systematic analysis of model behavior. We use Shared Interest to identify eight recurring patterns in model behavior, such as cases where contextual features or a subset of ground truth features are most important to the model. Working with representative real-world users, we show how Shared Interest can be used to decide if a model is trustworthy, uncover issues missed in manual analyses, and enable interactive probing.

📄 PDF Abstract BibTeX arXiv:2107.09234

Code (1)

mitvis/shared-interest 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

HOC 설명 없음

Similar Papers 제목 키워드 기반

Naturalistic measure of social norms alignment

2026-05-22 · Yevhen Kostiuk, Kenneth Enevoldsen, Peter Bjerregaard Vahlstrup, Márton Kardos 외 arxiv

Social norms reflect shared expectations on acceptable behavior. Measuring social norms alignment remains challenging, with existing approaches typically relying on artificial closed-form evaluations such as multiple-cho…

No Intruder, no Validity: Evaluation Criteria for Privacy-Preserving Text Anonymization

2021-03-16 · Maximilian Mozes, Bennett Kleinberg

For sensitive text data to be shared among NLP researchers and practitioners, shared documents need to comply with data protection and privacy laws. There is hence a growing interest in automated approaches for text anon…

AttributePrivacy PreservingText Anonymization

Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact

2026-03-01 · Michael Hardy, Yunsung Kim arxiv

LLMs increasingly excel on AI benchmarks, but doing so does not guarantee validity for downstream tasks. This study contrasts LLM alignment on benchmarks, downstream tasks, and, importantly the intended impact of those t…

Do Invariances in Deep Neural Networks Align with Human Perception?

2021-11-29 · Vedant Nanda, Ayan Majumdar, Camila Kolling, John P. Dickerson 외

An evaluation criterion for safe and trustworthy deep learning is how well the invariances captured by representations of deep neural networks (DNNs) are shared with humans. We identify challenges in measuring these inva…

Data AugmentationSelf-Supervised Learning

ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs

2024-09-15 · Hua Shen, Tiffany Knearem, Reshmi Ghosh, Yu-Ju Yang 외

As AI systems become more advanced, ensuring their alignment with a diverse range of individuals and societal values becomes increasingly critical. But how can we capture fundamental human values and assess the degree to…

Ethics