paper-with-me

Papers

Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models

2024-12-17 · Yingshui Tan, Boren Zheng, Baihui Zheng, Kerui Cao, Huiyun Jing, Jincheng Wei, Jiaheng Liu, Yancheng He, Wenbo Su, Xiangyong Zhu, Bo Zheng, Kaifu Zhang

With the rapid advancement of Large Language Models (LLMs), significant safety concerns have emerged. Fundamentally, the safety of large language models is closely linked to the accuracy, comprehensiveness, and clarity of their understanding of safety knowledge, particularly in domains such as law, policy and ethics. This factuality ability is crucial in determining whether these models can be deployed and applied safely and compliantly within specific regions. To address these challenges and better evaluate the factuality ability of LLMs to answer short questions, we introduce the Chinese SafetyQA benchmark. Chinese SafetyQA has several properties (i.e., Chinese, Diverse, High-quality, Static, Easy-to-evaluate, Safety-related, Harmless). Based on Chinese SafetyQA, we perform a comprehensive evaluation on the factuality abilities of existing LLMs and analyze how these capabilities relate to LLM abilities, e.g., RAG ability and robustness against attacks.

📄 PDF Abstract BibTeX arXiv:2412.15265

Code (0)

등록된 구현이 없습니다.

Tasks

EthicsFormRAG

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Adam 설명 없음
Weight Decay 설명 없음
Multi-Head Attention 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
WordPiece 설명 없음

Similar Papers 제목 키워드 기반

Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models

2024-11-11 · Yancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan 외

New LLM evaluation benchmarks are important to align with the rapid development of Large Language Models (LLMs). In this work, we present Chinese SimpleQA, the first comprehensive Chinese benchmark to evaluate the factua…

"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models

2025-02-17 · Jihao Gu, Yingyao Wang, Pi Bu, Chen Wang 외

The evaluation of factual accuracy in large vision language models (LVLMs) has lagged behind their rapid development, making it challenging to fully reflect these models' knowledge capacity and reliability. In this paper…

Object RecognitionQuestion AnsweringVisual Question Answering

The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation

2026-03-11 · Pavel Braslavski, Dmitrii Iarosh, Nikita Sushko, Andrey Sakhovskiy 외 arxiv

We present a configurable pipeline for generating multilingual sets of entities with specified characteristics, such as domain, geographical location and popularity, using data from Wikipedia and Wikidata. These datasets…

MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs

2025-10-27 · Yucheng Ning, Xixun Lin, Fang Fang, Yanan Cao arxiv

The widespread adoption of Large Language Models (LLMs) raises critical concerns about the factual accuracy of their outputs, especially in high-risk domains such as biomedicine, law, and education. Existing evaluation m…

Document-Level Event Factuality Identification via Adversarial Neural Network

2019-06-01 · NAACL 2019 6 · Zhong Qian, Peifeng Li, Qiaoming Zhu, Guodong Zhou

Document-level event factuality identification is an important subtask in event factuality and is crucial for discourse understanding in Natural Language Processing (NLP). Previous studies mainly suffer from the scarcity…

Sentence