paper-with-me

홈 › Papers

HateXScore: A Metric Suite for Evaluating Reasoning Quality in Hate Speech Explanations

2026-01-20 · Yujia Hu, Roy Ka-Wei Lee arxiv

Hateful speech detection is a key component of content moderation, yet current evaluation frameworks rarely assess why a text is deemed hateful. We introduce \textsf{HateXScore}, a four-component metric suite designed to evaluate the reasoning quality of model explanations. It assesses (i) conclusion explicitness, (ii) faithfulness and causal grounding of quoted spans, (iii) protected group identification (policy-configurable), and (iv) logical consistency among these elements. Evaluated on six diverse hate speech datasets, \textsf{HateXScore} is intended as a diagnostic complement to reveal interpretability failures and annotation inconsistencies that are invisible to standard metrics like Accuracy or F1. Moreover, human evaluation shows strong agreement with \textsf{HateXScore}, validating it as a practical tool for trustworthy and transparent moderation. \textcolor{red}{Disclaimer: This paper contains sensitive content that may be disturbing to some readers.}

📄 PDF Abstract BibTeX arXiv:2601.13547

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Neuro-Symbolic Benchmark Suite for Concept Quality and Reasoning Shortcuts

2024-06-14 · Samuele Bortolotti, Emanuele Marconato, Tommaso Carraro, Paolo Morettin 외

The advent of powerful neural classifiers has increased interest in problems that require both learning and reasoning. These problems are critical for understanding important properties of models, such as trustworthiness…

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

2025-02-13 · Dongzhi Jiang, Renrui Zhang, Ziyu Guo, Yanwei Li 외

Answering questions with Chain-of-Thought (CoT) has significantly enhanced the reasoning capabilities of Large Language Models (LLMs), yet its impact on Large Multimodal Models (LMMs) still lacks a systematic assessment …

BenchmarkingMathMMEMultimodal Reasoning+1

Evaluation Metric for Quality Control and Generative Models in Histopathology Images

2024-11-01 · Pranav Jeevan, Neeraj Nixon, Abhijeet Patil, Amit Sethi

Our study introduces ResNet-L2 (RL2), a novel metric for evaluating generative models and image quality in histopathology, addressing limitations of traditional metrics, such as Frechet inception distance (FID), when the…

Image Generation

Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks

2026-05-02 · Benjamin Warner, Ratna Sagari Grandhi, Max Kieffer, Aymane Ouraq 외 arxiv

Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insufficient coverage of relevant tasks. Existing suites have either satur…

Reinforcement LearningInformation ExtractionQuestion Answering

The Box is in the Pen: Evaluating Commonsense Reasoning in Neural Machine Translation

2025-03-05 · Findings of the Association for Computational Linguistics 2020 · Jie He, Tao Wang, Deyi Xiong, Qun Liu

Does neural machine translation yield translations that are congenial with common sense? In this paper, we present a test suite to evaluate the commonsense reasoning capability of neural machine translation. The test sui…

Common Sense ReasoningMachine TranslationSentenceTranslation