Min-Mid-Max Scaling, Limits of Agreement, and Agreement Score
In this paper, I solve a 60-year old question posed by Cohen's seminal paper (1960) and offer an agreement measure centered around the chance-expected agreement while isolating marginally forced agreement and disagreement. To achieve this, I formulate the minimum feasible agreement given row and column marginals by devising a new algorithm that minimizes the sum of diagonals in contingency tables. Based on this result, I also formulate the lower limit of the most common agreement measure-Cohen's kappa. Finally, I study the lower limit of maximum feasible agreement and devise two statistics of distribution similarity for agreement analysis.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
When to Vote, When to Rewrite: Disagreement-Guided Strategy Routing for Test-Time Scaling
Large Reasoning Models (LRMs) achieve strong performance on mathematical reasoning tasks but remain unreliable on challenging instances. Existing test-time scaling methods, such as repeated sampling, self-correction, and…
Mathematical ReasoningSelf-Agreement: A Framework for Fine-tuning Language Models to Find Agreement among Diverse Opinions
Finding an agreement among diverse opinions is a challenging topic in multiagent systems. Recently, large language models (LLMs) have shown great potential in addressing this challenge due to their remarkable capabilitie…
E-TCAV: Formalizing Penultimate Proxies for Efficient Concept Based Interpretability
TCAV (Testing with Concept Activation Vectors) is an interpretability method that assesses the alignment between the internal representations of a trained neural network and human-understandable, high-level concepts. Tho…
Decoupling Pixel Flipping and Occlusion Strategy for Consistent XAI Benchmarks
Feature removal is a central building block for eXplainable AI (XAI), both for occlusion-based explanations (Shapley values) as well as their evaluation (pixel flipping, PF). However, occlusion strategies can vary signif…
Your Pre-trained LLM is Secretly an Unsupervised Confidence Calibrator
Post-training of large language models is essential for adapting pre-trained language models (PLMs) to align with human preferences and downstream tasks. While PLMs typically exhibit well-calibrated confidence, post-trai…