paper-with-me

홈 › Papers

Towards Automatic Evaluation and Selection of PHI De-identification Models via Multi-Agent Collaboration

2025-10-17 · Guanchen Wu, Zuhui Chen, Yuzhang Xie, Carl Yang arxiv

Protected health information (PHI) de-identification is critical for enabling the safe reuse of clinical notes, yet evaluating and comparing PHI de-identification models typically depends on costly, small-scale expert annotations. We present TEAM-PHI, a multi-agent evaluation and selection framework that uses large language models (LLMs) to automatically measure de-identification quality and select the best-performing model without heavy reliance on gold labels. TEAM-PHI deploys multiple Evaluation Agents, each independently judging the correctness of PHI extractions and outputting structured metrics. Their results are then consolidated through an LLM-based majority voting mechanism that integrates diverse evaluator perspectives into a single, stable, and reproducible ranking. Experiments on a real-world clinical note corpus demonstrate that TEAM-PHI produces consistent and accurate rankings: despite variation across individual evaluators, LLM-based voting reliably converges on the same top-performing systems. Further comparison with ground-truth annotations and human evaluation confirms that the framework's automated rankings closely match supervised evaluation. By combining independent evaluation agents with LLM majority voting, TEAM-PHI offers a practical, secure, and cost-effective solution for automatic evaluation and best-model selection in PHI de-identification, even when ground-truth labels are limited.

📄 PDF Abstract BibTeX arXiv:2510.16194

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FrameNet automatic analysis : a study on a French corpus of encyclopedic texts

2018-12-19 · Gabriel Marzinotto, Géraldine Damnati, Frederic Bechet

This article presents an automatic frame analysis system evaluated on a corpus of French encyclopedic history texts annotated according to the FrameNet formalism. The chosen approach relies on an integrated sequence labe…

feature selection

Best Agent Identification for General Game Playing

2025-07-01 · Matthew Stephenson, Alex Newcombe, Eric Piette, Dennis Soemers arxiv

We present an efficient and generalised procedure to accurately identify the best (or near best) performing algorithm for each sub-task in a multi-problem domain. Our approach treats this as a set of best arm identificat…

Multi-Armed Bandits

Who Speaks Next? Multi-party AI Discussion Leveraging the Systematics of Turn-taking in Murder Mystery Games

2024-12-06 · Ryota Nonomura, Hiroki Mori

Multi-agent systems utilizing large language models (LLMs) have shown great promise in achieving natural dialogue. However, smooth dialogue control and autonomous decision making among agents still remain challenges. In …

Logical Reasoning

A Comparative User Evaluation of XRL Explanations using Goal Identification

2025-10-19 · Mark Towers, Yali Du, Christopher Freeman, Timothy J. Norman arxiv

Debugging is a core application of explainable reinforcement learning (XRL) algorithms; however, limited comparative evaluations have been conducted to understand their relative performance. We propose a novel evaluation…

Reinforcement Learning

AMIGO: Agentic Multi-Image Grounding Oracle Benchmark

2026-03-30 · Min Wang, Ata Mahjoubfar arxiv

Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness. We introduce AMIGO (Agentic Multi-Image Grounding Oracle Benchmark…

Question Selection