paper-with-me

홈 › Papers

CVC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models

2025-06-02 · Ping Wu, Guobin Shen, Dongcheng Zhao, Yuwei Wang, Yiting Dong, Yu Shi, Enmeng Lu, Feifei Zhao, Yi Zeng

Ensuring that Large Language Models (LLMs) align with mainstream human values and ethical norms is crucial for the safe and sustainable development of AI. Current value evaluation and alignment are constrained by Western cultural bias and incomplete domestic frameworks reliant on non-native rules; furthermore, the lack of scalable, rule-driven scenario generation methods makes evaluations costly and inadequate across diverse cultural contexts. To address these challenges, we propose a hierarchical value framework grounded in core Chinese values, encompassing three main dimensions, 12 core values, and 50 derived values. Based on this framework, we construct a large-scale Chinese Values Corpus (CVC) containing over 250,000 value rules enhanced and expanded through human annotation. Experimental results show that CVC-guided scenarios outperform direct generation ones in value boundaries and content diversity. In the evaluation across six sensitive themes (e.g., surrogacy, suicide), seven mainstream LLMs preferred CVC-generated options in over 70.5% of cases, while five Chinese human annotators showed an 87.5% alignment with CVC, confirming its universality, cultural relevance, and strong alignment with Chinese values. Additionally, we construct 400,000 rule-based moral dilemma scenarios that objectively capture nuanced distinctions in conflicting value prioritization across 17 LLMs. Our work establishes a culturally-adaptive benchmarking framework for comprehensive value evaluation and alignment, representing Chinese characteristics. All data are available at https://huggingface.co/datasets/Beijing-AISI/CVC, and the code is available at https://github.com/Beijing-AISI/CVC.

📄 PDF Abstract BibTeX arXiv:2506.01495

Code (1)

beijing-aisi/cvc 공식 구현

Tasks

Benchmarking

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

ChineseWebText: Large-scale High-quality Chinese Web Text Extracted with Effective Evaluation Model

2023-11-02 · Jianghao Chen, Pu Jian, Tengxiao Xi, Dongyi Yi 외

During the development of large language models (LLMs), the scale and quality of the pre-training data play a crucial role in shaping LLMs' capabilities. To accelerate the research of LLMs, several large-scale datasets, …

LSICC: A Large Scale Informal Chinese Corpus

2018-11-26 · Jianyu Zhao, Zhuoran Ji

Deep learning based natural language processing model is proven powerful, but need large-scale dataset. Due to the significant gap between the real-world tasks and existing Chinese corpus, in this paper, we introduce a l…

Chinese Word SegmentationDeep LearningSentiment Analysis

Development of a Web-Scale Chinese Word N-gram Corpus with Parts of Speech Information

2012-05-01 · LREC 2012 5 · Chi-Hsin Yu, Yi-jie Tang, Hsin-Hsi Chen

Web provides a large-scale corpus for researchers to study the language usages in real world. Developing a web-scale corpus needs not only a lot of computation resources, but also great efforts to handle the large variat…

Information RetrievalLanguage Modelling

CLUECorpus2020: A Large-scale Chinese Corpus for Pre-training Language Model

2020-03-03 · Liang Xu, Xuanwei Zhang, Qianqian Dong

In this paper, we introduce the Chinese corpus from CLUE organization, CLUECorpus2020, a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language gen…

8kLanguage ModelingLanguage ModellingSelf-Supervised Learning+1

A Large-Scale Chinese Short-Text Conversation Dataset

2020-08-10 · Yida Wang, Pei Ke, Yinhe Zheng, Kaili Huang 외

The advancements of neural dialogue generation models show promising results on modeling short-text conversations. However, training such models usually needs a large-scale high-quality dialogue corpus, which is hard to …

Dialogue GenerationShort-Text Conversation