paper-with-me

홈 › Papers

LLM Performance Predictors: Learning When to Escalate in Hybrid Human-AI Moderation Systems

2026-01-11 · Or Bachar, Or Levi, Sardhendu Mishra, Adi Levi, Manpreet Singh Minhas, Justin Miller, Omer Ben-Porat, Eilon Sheetrit, Jonathan Morra arxiv

As LLMs are increasingly integrated into human-in-the-loop content moderation systems, a central challenge is deciding when their outputs can be trusted versus when escalation for human review is preferable. We propose a novel framework for supervised LLM uncertainty quantification, learning a dedicated meta-model based on LLM Performance Predictors (LPPs) derived from LLM outputs: log-probabilities, entropy, and novel uncertainty attribution indicators. We demonstrate that our method enables cost-aware selective classification in real-world human-AI workflows: escalating high-risk cases while automating the rest. Experiments across state-of-the-art LLMs, including both off-the-shelf (Gemini, GPT) and open-source (Llama, Qwen), on multimodal and multilingual moderation tasks, show significant improvements over existing uncertainty estimators in accuracy-cost trade-offs. Beyond uncertainty estimation, the LPPs enhance explainability by providing new insights into failure conditions (e.g., ambiguous content vs. under-specified policy). This work establishes a principled framework for uncertainty-aware, scalable, and responsible human-AI moderation workflows.

📄 PDF Abstract BibTeX arXiv:2601.07006

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

2026-09-22 · Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman hf

LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify whe…

How angry are your customers? Sentiment analysis of support tickets that escalate

2020-10-26 · Colin Werner, Lloyd Montgomery, Sanja Dodos, Gabriel Tapuc 외

Software support ticket escalations can be an extremely costly burden for software organizations all over the world. Consequently, there exists an interest in researching how to better enable support analysts to handle s…

Sentiment Analysis

Generative Model-Enhanced Human Motion Prediction

2020-10-05 · Anthony Bourached, Ryan-Rhys Griffiths, Robert Gray, Ashwani Jha 외

The task of predicting human motion is complicated by the natural heterogeneity and compositionality of actions, necessitating robustness to distributional shifts as far as out-of-distribution (OoD). Here we formulate a …

Human motion predictionmodelmotion predictionPrediction

Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement

2024-07-25 · JaeHun Jung, Faeze Brahman, Yejin Choi

We present a principled approach to provide LLM-based evaluation with a rigorous guarantee of human agreement. We first propose that a reliable evaluation method should not uncritically rely on model preferences for pair…

Chatbot

DiagnoLLM: A Hybrid Bayesian Neural Language Framework for Interpretable Disease Diagnosis

2025-11-08 · Bowen Xu, Xinyue Zeng, Jiazhen Hu, Tuo Wang 외 arxiv

Building trustworthy clinical AI systems requires not only accurate predictions but also transparent, biologically grounded explanations. We present \texttt{DiagnoLLM}, a hybrid framework that integrates Bayesian deconvo…