paper-with-me

홈 › Papers

Libra: Large Chinese-based Safeguard for AI Content

2025-07-29 · Ziyang Chen, Huimu Yu, Xing Wu, Dongqin Liu, Songlin Hu arxiv

Large language models (LLMs) excel in text understanding and generation but raise significant safety and ethical concerns in high-stakes applications. To mitigate these risks, we present Libra-Guard, a cutting-edge safeguard system designed to enhance the safety of Chinese-based LLMs. Leveraging a two-stage curriculum training pipeline, Libra-Guard enhances data efficiency by employing guard pretraining on synthetic samples, followed by fine-tuning on high-quality, real-world data, thereby significantly reducing reliance on manual annotations. To enable rigorous safety evaluations, we also introduce Libra-Test, the first benchmark specifically designed to evaluate the effectiveness of safeguard systems for Chinese content. It covers seven critical harm scenarios and includes over 5,700 samples annotated by domain experts. Experiments show that Libra-Guard achieves 86.79% accuracy, outperforming Qwen2.5-14B-Instruct (74.33%) and ShieldLM-Qwen-14B-Chat (65.69%), and nearing closed-source models like Claude-3.5-Sonnet and GPT-4o. These contributions establish a robust framework for advancing the safety governance of Chinese LLMs and represent a tentative step toward developing safer, more reliable Chinese AI systems.

📄 PDF Abstract BibTeX arXiv:2507.21929

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards

2026-08-25 · Houcheng Jiang, Boxuan Zhang, Qiyong Zhong, Junfeng Fang 외 arxiv

Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting…

Reinforcement Learning

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models

2025-11-13 · Yuping Yan, Yuhan Xie, Yuanshuai Li, Yingchao Yu 외 arxiv

Since Multimodal Large Language Models (MLLMs) are increasingly being integrated into everyday tools and intelligent agents, growing concerns have arisen regarding their possible output of unsafe contents, ranging from t…

CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment

2026-06-13 · Wenbo Yu, Bohua Wang, Hao Fang, Kuofeng Gao 외 arxiv

Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns. While existing LLM safety guardrails excel in English or multilingual settings, they lack adaptation to C…

Prompt Engineering

Exploring the Privacy Protection Capabilities of Chinese Large Language Models

2024-03-27 · YuQi Yang, Xiaowen Huang, Jitao Sang

Large language models (LLMs), renowned for their impressive capabilities in various tasks, have significantly advanced artificial intelligence. Yet, these advancements have raised growing concerns about privacy and secur…

ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models

2024-10-24 · Hengxiang Zhang, Hongfu Gao, Qiang Hu, Guanhua Chen 외

With the rapid development of Large language models (LLMs), understanding the capabilities of LLMs in identifying unsafe content has become increasingly important. While previous works have introduced several benchmarks …