paper-with-me

Papers

Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection

2025-08-31 · Sanjeeevan Selvaganapathy, Mehwish Nasim arxiv

We investigate the efficacy of Large Language Models (LLMs) in detecting implicit and explicit hate speech, examining how models with minimal safety alignment (uncensored) compare with more heavily aligned (censored) counterparts in a deployed-model setting when deployed using political personas. While uncensored models are often framed as offering a less constrained perspective, our results reveal a trade-off: censored models outperform their uncensored counterparts in both accuracy and robustness, achieving 69.0\% versus 64.1\% strict accuracy. However, this higher performance is also associated with greater resistance to persona-based influence, while uncensored models are more malleable to ideological framing. Furthermore, we identify critical failures across all models in understanding nuanced language such as irony. We also find alarming fairness disparities in performance across different targeted groups and systemic overconfidence that renders self-reported certainty unreliable. These findings challenge the notion of LLMs as objective arbiters and highlight the need for more sophisticated auditing frameworks that account for fairness, calibration, and ideological consistency. Taken together, these results point to censorship-as-deployed rather than safety alignment in isolation as the more appropriate frame for interpreting model differences.

📄 PDF Abstract BibTeX arXiv:2509.00673

Code (0)

등록된 구현이 없습니다.

Tasks

Hate Speech Detection

Similar Papers 제목 키워드 기반

Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts

2025-11-25 · Xing Wang, Huiyuan Xie, Yiyan Wang, Chaojun Xiao 외 arxiv

Large language models (LLMs) are now deployed at unprecedented scale, assisting millions of users in daily tasks. However, the risk of these models assisting unlawful activities remains underexplored. In this study, we d…

For Those Who May Find Themselves on the Red Team

2025-11-23 · Tyler Shoemaker arxiv

This position paper argues that literary scholars must engage with large language model (LLM) interpretability research. While doing so will involve ideological struggle, if not out-right complicity, the necessity of thi…

ComplicitSplat: Downstream Models are Vulnerable to Blackbox Attacks by 3D Gaussian Splat Camouflages

2025-08-16 · Matthew Hull, Haoyang Yang, Pratham Mehta, Mansi Phute 외 arxiv

As 3D Gaussian Splatting (3DGS) gains rapid adoption in safety-critical tasks for efficient novel-view synthesis from static images, how might an adversary tamper images to cause harm? We introduce ComplicitSplat, the fi…

Calibrated and Efficient Sampling-Free Confidence Estimation for LiDAR Scene Semantic Segmentation

2024-11-18 · Hanieh Shojaei Miandashti, Qianqian Zou, Claus Brenner

Reliable deep learning models require not only accurate predictions but also well-calibrated confidence estimates to ensure dependable uncertainty estimation. This is crucial in safety-critical applications like autonomo…

Autonomous DrivingLIDAR Semantic SegmentationScene UnderstandingSegmentation+1

Safety-Aware Cascaded Inference for Crop Damage Assessment with Controlled Error Trade-offs

2026-07-28 · José Thiéry Messigbédé Hagbe, Gani Kawsar Gounou, Songbian Karim Zimé arxiv

In picture-based agricultural insurance for smallholder farmers, missed damage detections carry substantially higher cost than false alarms: a farmer who sustained real losses receives no payout, while unnecessary expert…