paper-with-me

홈 › Papers

Safety without alignment

2023-02-27 · András Kornai, Michael Bukatin, Zsolt Zombori

Currently, the dominant paradigm in AI safety is alignment with human values. Here we describe progress on developing an alternative approach to safety, based on ethical rationalism (Gewirth:1978), and propose an inherently safe implementation path via hybrid theorem provers in a sandbox. As AGIs evolve, their alignment may fade, but their rationality can only increase (otherwise more rational ones will have a significant evolutionary advantage) so an approach that ties their ethics to their rationality has clear long-term advantages.

📄 PDF Abstract BibTeX arXiv:2303.00752

Code (0)

등록된 구현이 없습니다.

Tasks

Ethics

Similar Papers 제목 키워드 기반

Jinx: Unlimited LLMs for Probing Alignment Failures

2025-08-11 · Jiahao Zhao, Liwei Dong arxiv

Unlimited, or so-called helpful-only language models are trained without safety alignment constraints and never refuse user queries. They are widely used by leading AI companies as internal tools for red teaming and alig…

Instruction FollowingRed Teaming

Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization

2025-12-12 · Yifan Niu, Han Xiao, Dongyi Liu, Nuo Chen 외 arxiv

As Large Language Models (LLMs) are increasingly deployed in real-world applications, it is important to ensure their behaviors align with human values, societal norms, and ethical principles. However, safety alignment u…

Reinforcement Learning

SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization

2025-11-17 · Xuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou 외 arxiv

Multimodal large language models (MLLMs) have demonstrated impressive reasoning and instruction-following capabilities, yet their expanded modality space introduces new compositional safety risks that emerge from complex…

MPO: Multilingual Safety Alignment via Reward Gap Optimization

2025-05-22 · Weixiang Zhao, Yulin Hu, Yang Deng, Tongtong Wu 외

Large language models (LLMs) have become increasingly central to AI applications worldwide, necessitating robust multilingual safety alignment to ensure secure deployment across diverse linguistic contexts. Existing pref…

Safety Alignment

Emerging Safety Attack and Defense in Federated Instruction Tuning of Large Language Models

2024-06-15 · Rui Ye, Jingyi Chai, Xiangrui Liu, Yaodong Yang 외

Federated learning (FL) enables multiple parties to collaboratively fine-tune an large language model (LLM) without the need of direct data sharing. Ideally, by training on decentralized data that is aligned with human p…

Federated LearningLanguage ModellingLarge Language ModelSafety Alignment