paper-with-me

Papers

TN-Eval: Rubric and Evaluation Protocols for Measuring the Quality of Behavioral Therapy Notes

2025-03-26 · Raj Sanjay Shah, Lei Xu, Qianchu Liu, Jon Burnsky, Drew Bertagnolli, Chaitanya Shivade

Behavioral therapy notes are important for both legal compliance and patient care. Unlike progress notes in physical health, quality standards for behavioral therapy notes remain underdeveloped. To address this gap, we collaborated with licensed therapists to design a comprehensive rubric for evaluating therapy notes across key dimensions: completeness, conciseness, and faithfulness. Further, we extend a public dataset of behavioral health conversations with therapist-written notes and LLM-generated notes, and apply our evaluation framework to measure their quality. We find that: (1) A rubric-based manual evaluation protocol offers more reliable and interpretable results than traditional Likert-scale annotations. (2) LLMs can mimic human evaluators in assessing completeness and conciseness but struggle with faithfulness. (3) Therapist-written notes often lack completeness and conciseness, while LLM-generated notes contain hallucination. Surprisingly, in a blind test, therapists prefer and judge LLM-generated notes to be superior to therapist-written notes.

📄 PDF Abstract BibTeX arXiv:2503.20648

Code (0)

등록된 구현이 없습니다.

Tasks

Hallucination

Similar Papers 제목 키워드 기반

LP-Eval: Rubric and Dataset for Measuring the Quality of Legal Proposition Generation

2026-05-19 · Shanshan Xu, Johan Lindholm, Amogh Raina, Henrik Palmer Olsen 외 arxiv

Legal proposition generation is central to legal reasoning and doctrinal scholarship, yet remain under-examined in Legal NLP. This paper investigates the automatic generation and evaluation of legal propositions from dec…

Legal Reasoning

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution

2026-05-08 · Mohit Raghavendra, Soham Dan, Miguel Romero Calvo, Yannis Yiming He 외 arxiv

We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test Writing (90 tasks), and Refactoring (70 tasks). SWE Atlas differs fro…

LLM-Driven Rubric-Based Assessment of Algebraic Competence in Multi-Stage Block Coding Tasks with Design and Field Evaluation

2025-10-04 · Yong Oh Lee, Byeonghun Bang, Sejun Oh arxiv

As online education platforms continue to expand, there is a growing need for assessment methods that not only measure answer accuracy but also capture the depth of students' cognitive processes in alignment with curricu…

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

2026-07-30 · Yuhang Zhu, Mingxuan Du, Benfeng Xu, Jie Gao 외 arxiv

Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliab…

ComplexConstraints and Beyond: Expert Rubrics for RLVR

2026-06-08 · Sushant Mehta, Liudas Panavas, Suhaas Garre, Edwin Chen arxiv

Evaluation protocols can lag behind LLM capabilities. Programmatically verified benchmarks cover narrow surface constraints, whereas real-world instruction following and agentic workflows require judging semantic, contex…

Instruction Following