paper-with-me

홈 › Papers

The Value of Context: Human versus Black Box Evaluators

2024-02-17 · Andrei Iakovlev, Annie Liang

Machine learning algorithms are now capable of performing evaluations previously conducted by human experts (e.g., medical diagnoses). How should we conceptualize the difference between evaluation by humans and by algorithms, and when should an individual prefer one over the other? We propose a framework to examine one key distinction between the two forms of evaluation: Machine learning algorithms are standardized, fixing a common set of covariates by which to assess all individuals, while human evaluators customize which covariates are acquired to each individual. Our framework defines and analyzes the advantage of this customization -- the value of context -- in environments with high-dimensional data. We show that unless the agent has precise knowledge about the joint distribution of covariates, the benefit of additional covariates generally outweighs the value of context.

📄 PDF Abstract BibTeX arXiv:2402.11157

Code (0)

등록된 구현이 없습니다.

Tasks

Medical Diagnosis

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Everyone prefers human writers, including AI

2025-10-09 · Wouter Haverals, Meredith Martin arxiv

As AI writing tools become widespread, we need to understand how both humans and machines evaluate literary style, a domain where objective standards are elusive and judgments are inherently subjective. We conducted cont…

Aligning Black-box Language Models with Human Judgments

2025-02-07 · Gerrit J. J. van den Burg, Gen Suzuki, Wei Liu, Murat Sensoy

Large language models (LLMs) are increasingly used as automated judges to evaluate recommendation systems, search engines, and other subjective tasks, where relying on human evaluators can be costly, time-consuming, and …

Recommendation Systems

Unveiling the Achilles' Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models

2024-05-23 · Yiming Chen, Chen Zhang, Danqing Luo, Luis Fernando D'Haro 외

The automatic evaluation of natural language generation (NLG) systems presents a long-lasting challenge. Recent studies have highlighted various neural metrics that align well with human evaluations. Yet, the robustness …

nlg evaluationText Generation

CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses

2024-07-15 · Jing Yao, Xiaoyuan Yi, Xing Xie

The rapid progress in Large Language Models (LLMs) poses potential risks such as generating unethical content. Assessing LLMs' values can help expose their misalignment, but relies on reference-free evaluators, e.g., fin…

EigenBench: A Comparative Behavioral Measure of Value Alignment

2025-09-02 · Jonathn Chang, Leonhard Piff, Suvadip Sana, Jasmine X. Li 외 arxiv

Aligning AI with human values is a pressing unsolved problem. To address the lack of quantitative metrics for value alignment, we propose EigenBench: a black-box method for comparatively benchmarking language models' val…