paper-with-me

Papers

Don't Command, Cultivate: An Exploratory Study of System-2 Alignment

2024-11-26 · Yuhang Wang, Yuxiang Zhang, Yanxu Zhu, Xinyan Wen, Jitao Sang

The o1 system card identifies the o1 models as the most robust within OpenAI, with their defining characteristic being the progression from rapid, intuitive thinking to slower, more deliberate reasoning. This observation motivated us to investigate the influence of System-2 thinking patterns on model safety. In our preliminary research, we conducted safety evaluations of the o1 model, including complex jailbreak attack scenarios using adversarial natural language prompts and mathematical encoding prompts. Our findings indicate that the o1 model demonstrates relatively improved safety performance; however, it still exhibits vulnerabilities, particularly against jailbreak attacks employing mathematical encoding. Through detailed case analysis, we identified specific patterns in the o1 model's responses. We also explored the alignment of System-2 safety in open-source models using prompt engineering and supervised fine-tuning techniques. Experimental results show that some simple methods to encourage the model to carefully scrutinize user requests are beneficial for model safety. Additionally, we proposed a implementation plan for process supervision to enhance safety alignment. The implementation details and experimental results will be provided in future versions.

📄 PDF Abstract BibTeX arXiv:2411.17075

Code (1)

adam-bjtu/system-2-alignment 공식 구현 pytorch

Tasks

Prompt EngineeringSafety Alignment

Similar Papers 제목 키워드 기반

Exploratory Experiments on Programming Autonomous Robots in Jadescript

2020-07-23 · Eleonora Iotti, Giuseppe Petrosino, Stefania Monica, Federico Bergenti

This paper describes exploratory experiments to validate the possibility of programming autonomous robots using an agent-oriented programming language. Proper perception of the environment, by means of various types of s…

Contrastive Language, Action, and State Pre-training for Robot Learning

2023-04-21 · Krishan Rana, Andrew Melnik, Niko Sünderhauf

In this paper, we introduce a method for unifying language, action, and state information in a shared embedding space to facilitate a range of downstream tasks in robot learning. Our method, Contrastive Language, Action,…

Retrieval

DWFF-Net : A Multi-Scale Farmland System Habitat Identification Method with Adaptive Dynamic Weight

2025-11-11 · Kesong Zheng, Zhi Song, Peizhou Li, Shuyi Yao 외 arxiv

Addressing the current lack of a standardized habitat classification system for cultivated land ecosystems, incomplete coverage of the habitat types, and the inability of existing models to effectively integrate semantic…

Culture in Action: Evaluating Text-to-Image Models through Social Activities

2025-11-07 · Sina Malakouti, Boqing Gong, Adriana Kovashka arxiv

Text-to-image (T2I) diffusion models achieve impressive photorealism by training on large-scale web data, but models inherit cultural biases and fail to depict underrepresented regions faithfully. Existing cultural bench…

Brain-Robot Interface for Exercise Mimicry

2025-09-14 · Carl Bettosi, Emilyann Nault, Lynne Baillie, Markus Garschall 외 arxiv

For social robots to maintain long-term engagement as exercise instructors, rapport-building is essential. Motor mimicry--imitating one's physical actions--during social interaction has long been recognized as a powerful…