paper-with-me

홈 › Papers

Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection

2024-02-18 · Min Zhang, Jianfeng He, Taoran Ji, Chang-Tien Lu

The fairness and trustworthiness of Large Language Models (LLMs) are receiving increasing attention. Implicit hate speech, which employs indirect language to convey hateful intentions, occupies a significant portion of practice. However, the extent to which LLMs effectively address this issue remains insufficiently examined. This paper delves into the capability of LLMs to detect implicit hate speech (Classification Task) and express confidence in their responses (Calibration Task). Our evaluation meticulously considers various prompt patterns and mainstream uncertainty estimation methods. Our findings highlight that LLMs exhibit two extremes: (1) LLMs display excessive sensitivity towards groups or topics that may cause fairness issues, resulting in misclassifying benign statements as hate speech. (2) LLMs' confidence scores for each method excessively concentrate on a fixed range, remaining unchanged regardless of the dataset's complexity. Consequently, the calibration performance is heavily reliant on primary classification accuracy. These discoveries unveil new limitations of LLMs, underscoring the need for caution when optimizing models to ensure they do not veer towards extremes. This serves as a reminder to carefully consider sensitivity and confidence in the pursuit of model fairness.

📄 PDF Abstract BibTeX arXiv:2402.11406

Code (0)

등록된 구현이 없습니다.

Tasks

FairnessHate Speech DetectionSensitivity

Similar Papers 제목 키워드 기반

Validation of ML-UQ calibration statistics using simulated reference values: a sensitivity analysis

2024-03-01 · Pascal Pernot

Some popular Machine Learning Uncertainty Quantification (ML-UQ) calibration statistics do not have predefined reference values and are mostly used in comparative studies. In consequence, calibration is almost never vali…

DiagnosticSensitivityUncertainty Quantification

Machine Learning based Parameter Sensitivity of Regional Climate Models -- A Case Study of the WRF Model for Heat Extremes over Southeast Australia

2023-07-27 · P. Jyoteeshkumar Reddy, Sandeep Chinta, Richard Matear, John Taylor 외

Heatwaves and bushfires cause substantial impacts on society and ecosystems across the globe. Accurate information of heat extremes is needed to support the development of actionable mitigation and adaptation strategies.…

Sensitivity

Modelling and simulating spatial extremes by combining extreme value theory with generative adversarial networks

2021-10-30 · Younes Boulaguiem, Jakob Zscheischler, Edoardo Vignotto, Karin van der Wiel 외

Modelling dependencies between climate extremes is important for climate risk assessment, for instance when allocating emergency management funds. In statistics, multivariate extreme value theory is often used to model s…

Management

A Market-Clearing-based Sensitivity Model for Locational Marginal and Average Carbon Emission

2024-01-24 · Zelong Lu

This letter proposes a market-clearing-based locational marginal carbon emission (LMCE) metric to assess the marginal carbon emission effect of nodal load demand. Unlike the prevalent carbon emission flow (CEF) method th…

Sensitivity

Guarding the Meaning: Self-Supervised Training for Semantic Robustness in Guard Models

2025-11-06 · Cristina Pinneri, Christos Louizos arxiv

Guard models are a critical component of LLM safety, but their sensitivity to superficial linguistic variations remains a key vulnerability. We show that even meaning-preserving paraphrases can cause large fluctuations i…