paper-with-me

홈 › Papers

A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

2026-05-28 · Kenji Imamura, Masao Ideuchi, Atsushi Fujita arxiv

In this paper, we discuss question-answer dataset for LLM safety evaluation, with a focus on illegal activities. Specifically, on the basis of manual analysis of AnswerCarefully, we introduce several additional information, methods for creating question-answer examples, and a rubric for evaluating LLM-generated responses. The outcomes of this study are intended to be shared with the "JAI-Trust" project.

📄 PDF Abstract BibTeX arXiv:2605.29340

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models

2024-06-14 · Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu, Meijuan An 외

With the profound development of large language models(LLMs), their safety concerns have garnered increasing attention. However, there is a scarcity of Chinese safety benchmarks for LLMs, and the existing safety taxonomi…

Multiple-choiceQuestion Answering

Quality of Answers of Generative Large Language Models vs Peer Patients for Interpreting Lab Test Results for Lay Patients: Evaluation Study

2024-01-23 · Zhe He, Balu Bhasuran, Qiao Jin, Shubo Tian 외

Lab results are often confusing and hard to understand. Large language models (LLMs) such as ChatGPT have opened a promising avenue for patients to get their questions answered. We aim to assess the feasibility of using …

NewsBench: A Systematic Evaluation Framework for Assessing Editorial Capabilities of Large Language Models in Chinese Journalism

2024-02-29 · Miao Li, Ming-Bin Chen, Bo Tang, Shengbin Hou 외

We present NewsBench, a novel evaluation framework to systematically assess the capabilities of Large Language Models (LLMs) for editorial capabilities in Chinese journalism. Our constructed benchmark dataset is focused …

EthicsMultiple-choice

Fake Alignment: Are LLMs Really Aligned Well?

2023-11-10 · Yixu Wang, Yan Teng, Kexin Huang, Chengqi Lyu 외

The growing awareness of safety concerns in large language models (LLMs) has sparked considerable interest in the evaluation of safety. This study investigates an under-explored issue about the evaluation of LLMs, namely…

Multiple-choice

Memory-Augmented Knowledge Fusion with Safety-Aware Decoding for Domain-Adaptive Question Answering

2025-12-02 · Lei Fu, Xiang Chen, Kaige Gao Xinyue Huang, Kejian Tong arxiv

Domain-specific question answering (QA) systems for services face unique challenges in integrating heterogeneous knowledge sources while ensuring both accuracy and safety. Existing large language models often struggle wi…

Question Answering