paper-with-me

Papers

Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance

2024-07-10 · Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, Nouha Dziri, Dan Jurafsky, Maarten Sap

The ability to communicate uncertainty, risk, and limitation is crucial for the safety of large language models. However, current evaluations of these abilities rely on simple calibration, asking whether the language generated by the model matches appropriate probabilities. Instead, evaluation of this aspect of LLM communication should focus on the behaviors of their human interlocutors: how much do they rely on what the LLM says? Here we introduce an interaction-centered evaluation framework called Rel-A.I. (pronounced "rely"}) that measures whether humans rely on LLM generations. We use this framework to study how reliance is affected by contextual features of the interaction (e.g, the knowledge domain that is being discussed), or the use of greetings communicating warmth or competence (e.g., "I'm happy to help!"). We find that contextual characteristics significantly affect human reliance behavior. For example, people rely 10% more on LMs when responding to questions involving calculations and rely 30% more on LMs that are perceived as more competent. Our results show that calibration and language quality alone are insufficient in evaluating the risks of human-LM interactions, and illustrate the need to consider features of the interactional context.

📄 PDF Abstract BibTeX arXiv:2407.07950

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Towards interactive evaluations for interaction harms in human-AI systems

2024-05-17 · Lujain Ibrahim, Saffron Huang, Umang Bhatt, Lama Ahmad 외

Current AI evaluation paradigms that rely on static, model-only tests fail to capture harms that emerge through sustained human-AI interaction. As interactive AI systems, such as AI companions, proliferate in daily life,…

Ethics

Measuring and mitigating overreliance to build human-compatible AI

2025-09-08 · Lujain Ibrahim, Katherine M. Collins, Sunnie S. Y. Kim, Anka Reuel 외 arxiv

Large language models (LLMs) distinguish themselves from previous technologies by functioning as collaborative ``thought partners,'' capable of engaging more fluidly in natural language on a range of tasks. As LLMs incre…

From Accuracy to Readiness: Metrics and Benchmarks for Human-AI Decision-Making

2026-03-19 · Min Hun Lee arxiv

Artificial intelligence (AI) systems are deployed as collaborators in human decision-making. Yet, evaluation practices focus primarily on model accuracy rather than whether human-AI teams are prepared to collaborate safe…

A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice

2026-06-04 · Ranjan Mishra, Jakob Schoeffer arxiv

Appropriate reliance on AI advice has become a central research theme in human-AI collaboration. Existing frameworks have focused exclusively on point predictions as AI advice. However, set-valued AI advice (e.g., discre…

Decision Making

A Survey of AI Reliance

2024-07-22 · Sven Eckhardt, Niklas Kühl, Mateusz Dolata, Gerhard Schwabe

Artificial intelligence (AI) systems have become an indispensable component of modern technology. However, research on human behavioral responses is lagging behind, i.e., the research into human reliance on AI advice (AI…

Survey