paper-with-me

홈 › Papers

FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity

2023-11-30 · Shiyao Cui, Zhenyu Zhang, Yilong Chen, Wenyuan Zhang, Tianyun Liu, Siqi Wang, Tingwen Liu

The widespread of generative artificial intelligence has heightened concerns about the potential harms posed by AI-generated texts, primarily stemming from factoid, unfair, and toxic content. Previous researchers have invested much effort in assessing the harmlessness of generative language models. However, existing benchmarks are struggling in the era of large language models (LLMs), due to the stronger language generation and instruction following capabilities, as well as wider applications. In this paper, we propose FFT, a new benchmark with 2116 elaborated-designed instances, for LLM harmlessness evaluation with factuality, fairness, and toxicity. To investigate the potential harms of LLMs, we evaluate 9 representative LLMs covering various parameter scales, training stages, and creators. Experiments show that the harmlessness of LLMs is still under-satisfactory, and extensive analysis derives some insightful findings that could inspire future research for harmless LLM research.

📄 PDF Abstract BibTeX arXiv:2311.18580

Code (1)

cuishiyao96/fft 공식 구현

Tasks

FairnessInstruction FollowingText Generation

Similar Papers 제목 키워드 기반

Flames: Benchmarking Value Alignment of LLMs in Chinese

2023-11-12 · Kexin Huang, Xiangyang Liu, Qianyu Guo, Tianxiang Sun 외

The widespread adoption of large language models (LLMs) across various regions underscores the urgent need to evaluate their alignment with human values. Current benchmarks, however, fall short of effectively uncovering …

BenchmarkingFairness

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures

2025-06-09 · Yukai Zhou, Sibei Yang, Wenjie Wang

Large language models (LLMs) are increasingly deployed in real-world applications, raising concerns about their security. While jailbreak attacks highlight failures under overtly harmful queries, they overlook a critical…

When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation

2025-10-08 · Xunyi Jiang, Dingyi Chang, Julian McAuley, Xin Xu arxiv

The rapid evolution of large language models (LLMs) and the real world has outpaced the static nature of widely used evaluation benchmarks, raising concerns about their reliability for evaluating LLM factuality. While su…

BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models

2025-03-31 · Alok Abhishek, Lisa Erickson, Tushar Bandopadhyay

In this research, we introduce BEATS, a novel framework for evaluating Bias, Ethics, Fairness, and Factuality in Large Language Models (LLMs). Building upon the BEATS framework, we present a bias benchmark for LLMs that …

EthicsFairnessMisinformation

Global-Liar: Factuality of LLMs over Time and Geographic Regions

2024-01-31 · Shujaat Mirza, Bruno Coelho, Yuyuan Cui, Christina Pöpper 외

The increasing reliance on AI-driven solutions, particularly Large Language Models (LLMs) like the GPT series, for information retrieval highlights the critical need for their factuality and fairness, especially amidst t…

FairnessInformation RetrievalMisinformation