paper-with-me

홈 › Papers

Toward Responsible Federated Large Language Models: Leveraging a Safety Filter and Constitutional AI

2025-02-23 · Eunchung Noh, Jeonghun Baek

Recent research has increasingly focused on training large language models (LLMs) using federated learning, known as FedLLM. However, responsible AI (RAI), which aims to ensure safe responses, remains underexplored in the context of FedLLM. In FedLLM, client data used for training may contain harmful content, leading to unsafe LLMs that generate harmful responses. Aggregating such unsafe LLMs into the global model and distributing them to clients may result in the widespread deployment of unsafe LLMs. To address this issue, we incorporate two well-known RAI methods into FedLLM: the safety filter and constitutional AI. Our experiments demonstrate that these methods significantly enhance the safety of the LLM, achieving over a 20% improvement on AdvBench, a benchmark for evaluating safety performance.

📄 PDF Abstract BibTeX arXiv:2502.16691

Code (0)

등록된 구현이 없습니다.

Tasks

Federated Learning

Similar Papers 제목 키워드 기반

FairSense-AI: Responsible AI Meets Sustainability

2025-03-04 · Shaina Raza, Mukund Sayeeganesh Chettiar, Matin Yousefabadi, Tahniat Khan 외

In this paper, we introduce FairSense-AI: a multimodal framework designed to detect and mitigate bias in both text and images. By leveraging Large Language Models (LLMs) and Vision-Language Models (VLMs), FairSense-AI un…

Fairness

NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models

2025-04-29 · Yi Zhou, Wenpeng Xing, Dezhang Kong, Changting Lin 외

Safety alignment in large language models (LLMs) is achieved through fine-tuning mechanisms that regulate neuron activations to suppress harmful content. In this work, we propose a novel approach to induce disalignment b…

Safety Alignment

ResponsibleRobotBench: Benchmarking Responsible Robot Manipulation using Multi-modal Large Language Models

2025-12-03 · Lei Zhang, Ju Dong, Kaixin Bai, Minheng Ni 외 arxiv

Recent advances in large multimodal models have enabled new opportunities in embodied AI, particularly in robotic manipulation. These models have shown strong potential in generalization and reasoning, but achieving reli…

Robot Manipulation

Emerging Safety Attack and Defense in Federated Instruction Tuning of Large Language Models

2024-06-15 · Rui Ye, Jingyi Chai, Xiangrui Liu, Yaodong Yang 외

Federated learning (FL) enables multiple parties to collaboratively fine-tune an large language model (LLM) without the need of direct data sharing. Ideally, by training on decentralized data that is aligned with human p…

Federated LearningLanguage ModellingLarge Language ModelSafety Alignment

Responsible AI in Construction Safety: Systematic Evaluation of Large Language Models and Prompt Engineering

2024-11-13 · Farouq Sammour, Jia Xu, Xi Wang, Mo Hu 외

Construction remains one of the most hazardous sectors. Recent advancements in AI, particularly Large Language Models (LLMs), offer promising opportunities for enhancing workplace safety. However, responsible integration…

ManagementPrompt Engineering