paper-with-me

Papers

Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories

2025-03-28 · Yazhou Zhang, Qimeng Liu, Qiuchi Li, Peng Zhang, Jing Qin

Evaluating the value alignment of large language models (LLMs) has traditionally relied on single-sentence adversarial prompts, which directly probe models with ethically sensitive or controversial questions. However, with the rapid advancements in AI safety techniques, models have become increasingly adept at circumventing these straightforward tests, limiting their effectiveness in revealing underlying biases and ethical stances. To address this limitation, we propose an upgraded value alignment benchmark that moves beyond single-sentence prompts by incorporating multi-turn dialogues and narrative-based scenarios. This approach enhances the stealth and adversarial nature of the evaluation, making it more robust against superficial safeguards implemented in modern LLMs. We design and implement a dataset that includes conversational traps and ethically ambiguous storytelling, systematically assessing LLMs' responses in more nuanced and context-rich settings. Experimental results demonstrate that this enhanced methodology can effectively expose latent biases that remain undetected in traditional single-shot evaluations. Our findings highlight the necessity of contextual and dynamic testing for value alignment in LLMs, paving the way for more sophisticated and realistic assessments of AI ethics and safety.

📄 PDF Abstract BibTeX arXiv:2503.22115

Code (0)

등록된 구현이 없습니다.

Tasks

EthicsSentence

Similar Papers 제목 키워드 기반

Is ChatGPT a Good Causal Reasoner? A Comprehensive Evaluation

2023-05-12 · Jinglong Gao, Xiao Ding, Bing Qin, Ting Liu

Causal reasoning ability is crucial for numerous NLP applications. Despite the impressive emerging ability of ChatGPT in various NLP tasks, it is unclear how well ChatGPT performs in causal reasoning. In this paper, we c…

HallucinationIn-Context Learning

Learning and Upgrading in Global Value Chains: An Analysis of India's Manufacturing Sector

2021-01-12 · Sourish Dutta

The topic of my research is "Learning and Upgrading in Global Value Chains: An Analysis of India's Manufacturing Sector". To analyse India's learning and upgrading through position, functions, specialisation & value addi…

Position

D2CSE: Difference-aware Deep continuous prompts for Contrastive Sentence Embeddings

2023-04-18 · Hyunjae Lee

This paper describes Difference-aware Deep continuous prompt for Contrastive Sentence Embeddings (D2CSE) that learns sentence embeddings. Compared to state-of-the-art approaches, D2CSE computes sentence vectors that are …

Contrastive LearningRetrievalSemantic Textual SimilaritySentence+2

VideoBooth: Diffusion-based Video Generation with Image Prompts

2023-12-01 · CVPR 2024 1 · Yuming Jiang, Tianxing Wu, Shuai Yang, Chenyang Si 외

Text-driven video generation witnesses rapid progress. However, merely using text prompts is not enough to depict the desired subject appearance that accurately aligns with users' intents, especially for customized conte…

Video Generation

Multi-Sentence Grounding for Long-term Instructional Video

2023-12-21 · Zeqian Li, Qirui Chen, Tengda Han, Ya zhang 외

In this paper, we aim to establish an automatic, scalable pipeline for denoising the large-scale instructional dataset and construct a high-quality video-text dataset with multiple descriptive steps supervision, named Ho…

DenoisingDescriptiveLanguage ModellingLarge Language Model+3