paper-with-me

홈 › Papers

Measuring Error Alignment for Decision-Making Systems

2024-09-20 · Binxia Xu, Antonis Bikakis, Daniel Onah, Andreas Vlachidis, Luke Dickens

Given that AI systems are set to play a pivotal role in future decision-making processes, their trustworthiness and reliability are of critical concern. Due to their scale and complexity, modern AI systems resist direct interpretation, and alternative ways are needed to establish trust in those systems, and determine how well they align with human values. We argue that good measures of the information processing similarities between AI and humans, may be able to achieve these same ends. While Representational alignment (RA) approaches measure similarity between the internal states of two systems, the associated data can be expensive and difficult to collect for human systems. In contrast, Behavioural alignment (BA) comparisons are cheaper and easier, but questions remain as to their sensitivity and reliability. We propose two new behavioural alignment metrics misclassification agreement which measures the similarity between the errors of two systems on the same instances, and class-level error similarity which measures the similarity between the error distributions of two systems. We show that our metrics correlate well with RA metrics, and provide complementary information to another BA metric, within a range of domains, and set the scene for a new approach to value alignment.

📄 PDF Abstract BibTeX arXiv:2409.13919

Code (1)

xubinxia/error_align 공식 구현

Tasks

Decision Making

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language Models

2024-10-03 · Yinhong Liu, Zhijiang Guo, Tianya Liang, Ehsan Shareghi 외

Large Language Models (LLMs) are expected to be predictable and trustworthy to support reliable decision-making systems. Yet current LLMs often show inconsistencies in their judgments. In this work, we examine logical pr…

Decision MakingNegation

Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency

2020-06-30 · NeurIPS 2020 12 · Robert Geirhos, Kristof Meding, Felix A. Wichmann

A central problem in cognitive science and behavioural neuroscience as well as in machine learning and artificial intelligence research is to ascertain whether two or more decision makers (be they brains or algorithms) u…

Decision MakingObject Recognition

Making Large Language Models Better Planners with Reasoning-Decision Alignment

2024-08-25 · Zhijian Huang, Tao Tang, Shaoxiang Chen, Sihao Lin 외

Data-driven approaches for autonomous driving (AD) have been widely adopted in the past decade but are confronted with dataset bias and uninterpretability. Inspired by the knowledge-driven nature of human driving, recent…

Autonomous DrivingDecision MakingScene Understanding

Turing Representational Similarity Analysis (RSA): A Flexible Method for Measuring Alignment Between Human and Artificial Intelligence

2024-11-30 · Mattson Ogg, Ritwik Bose, Jamie Scharf, Christopher Ratto 외

As we consider entrusting Large Language Models (LLMs) with key societal and decision-making roles, measuring their alignment with human cognition becomes critical. This requires methods that can assess how these systems…

Decision-Theoretic Safety Assessment of Persona-Driven Multi-Agent Systems in O-RAN

2026-04-03 · Zeinab Nezami, Syed Ali Raza Zaidi, Maryam Hafeez, Louis Powell 외 arxiv

Autonomous network management in Open Radio Access Networks requires intelligent decision making across conflicting objectives, yet existing LLM based multi agent systems employ homogeneous strategies and lack systematic…

Code GenerationDecision Making