paper-with-me

홈 › Papers

LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data

2025-08-16 · Stephen Meisenbacher, Alexandra Klymenko, Florian Matthes arxiv

Despite advances in the field of privacy-preserving Natural Language Processing (NLP), a significant challenge remains the accurate evaluation of privacy. As a potential solution, using LLMs as a privacy evaluator presents a promising approach $\unicode{x2013}$ a strategy inspired by its success in other subfields of NLP. In particular, the so-called $\textit{LLM-as-a-Judge}$ paradigm has achieved impressive results on a variety of natural language evaluation tasks, demonstrating high agreement rates with human annotators. Recognizing that privacy is both subjective and difficult to define, we investigate whether LLM-as-a-Judge can also be leveraged to evaluate the privacy sensitivity of textual data. Furthermore, we measure how closely LLM evaluations align with human perceptions of privacy in text. Resulting from a study involving 10 datasets, 13 LLMs, and 677 human survey participants, we confirm that privacy is indeed a difficult concept to measure empirically, exhibited by generally low inter-human agreement rates. Nevertheless, we find that LLMs can accurately model a global human privacy perspective, and through an analysis of human and LLM reasoning patterns, we discuss the merits and limitations of LLM-as-a-Judge for privacy evaluation in textual data. Our findings pave the way for exploring the feasibility of LLMs as privacy evaluators, addressing a core challenge in solving pressing privacy issues with innovative technical solutions.

📄 PDF Abstract BibTeX arXiv:2508.12158

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PrivacyAlign: Contextual Privacy Alignment for LLM Agents

2026-06-19 · Manveer Singh Tamber, Abhay Puri, Marc-Etienne Brunet, Perouz Taslakian 외 arxiv

AI agents acting on behalf of users are constantly making decisions, and for users to trust their agents, those decisions must align with what they actually want. Privacy is an important alignment problem for agents: eve…

Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

2024-08-23 · Hui Wei, Shenghua He, Tian Xia, Fei Liu 외

LLM-as-a-Judge has been widely applied to evaluate and compare different LLM alignmnet approaches (e.g., RLHF and DPO). However, concerns regarding its reliability have emerged, due to LLM judges' biases and inconsistent…

Through the Lens of Split Vote: Exploring Disagreement, Difficulty and Calibration in Legal Case Outcome Classification

2024-02-11 · Shanshan Xu, T. Y. S. S Santosh, Oana Ichim, Barbara Plank 외

In legal decisions, split votes (SV) occur when judges cannot reach a unanimous decision, posing a difficulty for lawyers who must navigate diverse legal arguments and opinions. In high-stakes domains, understanding the …

Navigate

Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment

2026-08-21 · Emma Granqvist, Rocío Mercado, Samuel Genheden arxiv

Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. Reference-based met…

Drug Discovery

Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions

2024-08-16 · Bhuvanashree Murugadoss, Christian Poelitz, Ian Drosos, Vu Le 외

LLMs-as-a-judge is a recently popularized method which replaces human judgements in task evaluation (Zheng et al. 2024) with automatic evaluation using LLMs. Due to widespread use of RLHF (Reinforcement Learning from Hum…