paper-with-me

Papers

Evaluating Human-Language Model Interaction

2022-12-19 · Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, Rose E. Wang, Minae Kwon, Joon Sung Park, Hancheng Cao, Tony Lee, Rishi Bommasani, Michael Bernstein, Percy Liang

Many real-world applications of language models (LMs), such as writing assistance and code autocomplete, involve human-LM interaction. However, most benchmarks are non-interactive in that a model produces output without human involvement. To evaluate human-LM interaction, we develop a new framework, Human-AI Language-based Interaction Evaluation (HALIE), that defines the components of interactive systems and dimensions to consider when designing evaluation metrics. Compared to standard, non-interactive evaluation, HALIE captures (i) the interactive process, not only the final output; (ii) the first-person subjective experience, not just a third-party assessment; and (iii) notions of preference beyond quality (e.g., enjoyment and ownership). We then design five tasks to cover different forms of interaction: social dialogue, question answering, crossword puzzles, summarization, and metaphor generation. With four state-of-the-art LMs (three variants of OpenAI's GPT-3 and AI21 Labs' Jurassic-1), we find that better non-interactive performance does not always translate to better human-LM interaction. In particular, we highlight three cases where the results from non-interactive and interactive metrics diverge and underscore the importance of human-LM interaction for LM evaluation.

📄 PDF Abstract BibTeX arXiv:2212.09746

Code (1)

stanford-crfm/halie 공식 구현

Tasks

Language ModelingLanguage ModellingmodelQuestion Answering

Methods 이 논문이 사용한 방법론

{Dispute@FaQ-s}How to file a dispute with Expedia? How to file a dispute with Expedia? To file a complaint against Expedia, first try contacting their customer service directly. You can reach them by phone at…
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Weight Decay 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음

Similar Papers 제목 키워드 기반

Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation

2025-09-03 · Reina Ishikawa, Ryo Fujii, Hideo Saito, Ryo Hachiuma arxiv

Evaluating concept customization is challenging, as it requires a comprehensive assessment of fidelity to generative prompts and concept images. Moreover, evaluating multiple concepts is considerably more difficult than …

ProSPer: Probing Human and Neural Network Language Model Understanding of Spatial Perspective

2021-11-01 · EMNLP (BlackboxNLP) 2021 11 · Tessa Masis, Carolyn Anderson

Understanding perspectival language is important for applications like dialogue systems and human-robot interaction. We propose a probe task that explores how well language models understand spatial perspective. We prese…

Language ModelingLanguage Modelling

Towards interactive evaluations for interaction harms in human-AI systems

2024-05-17 · Lujain Ibrahim, Saffron Huang, Umang Bhatt, Lama Ahmad 외

Current AI evaluation paradigms that rely on static, model-only tests fail to capture harms that emerge through sustained human-AI interaction. As interactive AI systems, such as AI companions, proliferate in daily life,…

Ethics

Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance

2024-07-10 · Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, Nouha Dziri 외

The ability to communicate uncertainty, risk, and limitation is crucial for the safety of large language models. However, current evaluations of these abilities rely on simple calibration, asking whether the language gen…

Sentence

V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction

2025-03-22 · Yiming Zhao, Yu Zeng, Yukun Qi, Yaoyang Liu 외

Large Vision-Language Models (LVLMs) have made significant progress in the field of video understanding recently. However, current benchmarks uniformly lean on text prompts for evaluation, which often necessitate complex…

BenchmarkingVideo Understanding