paper-with-me

Papers

SafeLM: Unified Privacy-Aware Optimization for Trustworthy Federated Large Language Models

2026-04-17 · Noor Islam S. Mohammad, Uluğ Bayazıt arxiv

Large language models (LLMs) are increasingly deployed in high-stakes domains, yet a unified treatment of their overlapping safety challenges remains lacking. We present SafeLM, a framework that jointly addresses four pillars of LLM safety: privacy, security, misinformation, and adversarial robustness. SafeLM combines federated training with gradient smartification and Paillier encryption for privacy, integrates defenses against training and inference-time attacks, employs contrastive grounding with calibrated decoding to reduce hallucinations, and introduces alignment-aware binarized aggregation to enhance robustness while maintaining bounded reconstruction quality. Across benchmarks on factuality, toxicity, and membership inference, SafeLM achieves 98.0% harmful content detection accuracy, reduces communication by 96.9%, and lowers gradient inversion PSNR from 31.7 dB to 15.1 dB. Ablations show that each component contributes independently, whereas their integration yields a strong privacy utility efficiency trade-off for deploying trustworthy LLMs.

📄 PDF Abstract BibTeX arXiv:2604.16606

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial Robustness

Similar Papers 제목 키워드 기반

Information Theoretic Evaluation of Privacy-Leakage, Interpretability, and Transferability for Trustworthy AI

2021-06-06 · Mohit Kumar, Bernhard A. Moser, Lukas Fischer, Bernhard Freudenthaler

In order to develop machine learning and deep learning models that take into account the guidelines and principles of trustworthy AI, a novel information theoretic trustworthy AI framework is introduced. A unified approa…

Heart Rate VariabilityPrivacy Preserving

Unraveling Privacy Risks of Individual Fairness in Graph Neural Networks

2023-01-30 · He Zhang, Xingliang Yuan, Shirui Pan

Graph neural networks (GNNs) have gained significant attraction due to their expansive real-world applications. To build trustworthy GNNs, two aspects - fairness and privacy - have emerged as critical considerations. Pre…

Fairness

Privacy and Transparency in Graph Machine Learning: A Unified Perspective

2022-07-22 · Megha Khosla

Graph Machine Learning (GraphML), whereby classical machine learning is generalized to irregular graph domains, has enjoyed a recent renaissance, leading to a dizzying array of models and their applications in several do…

BIG-bench Machine LearningGraph LearningPosition

Trustworthy Distributed AI Systems: Robustness, Privacy, and Governance

2024-02-02 · Wenqi Wei, Ling Liu

Emerging Distributed AI systems are revolutionizing big data computing and data processing capabilities with growing economic and societal impact. However, recent studies have identified new attack surfaces and risks cau…

Fairness

Trustworthy Quantum Machine Learning: A Roadmap for Reliability, Robustness, and Security in the NISQ Era

2025-11-04 · Ferhat Ozgur Catak, Jungwon Seo, Umit Cali arxiv

Quantum machine learning (QML) is a promising paradigm for tackling computational problems that challenge classical AI. Yet, the inherent probabilistic behavior of quantum mechanics, device noise in NISQ hardware, and hy…

Quantum Machine LearningAdversarial RobustnessDecision Making