paper-with-me

Papers

A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy

2025-01-16 · Huandong Wang, Wenjie Fu, Yingzhou Tang, Zhilong Chen, Yuxi Huang, Jinghua Piao, Chen Gao, Fengli Xu, Tao Jiang, Yong Li

While large language models (LLMs) present significant potential for supporting numerous real-world applications and delivering positive social impacts, they still face significant challenges in terms of the inherent risk of privacy leakage, hallucinated outputs, and value misalignment, and can be maliciously used for generating toxic content and unethical purposes after been jailbroken. Therefore, in this survey, we present a comprehensive review of recent advancements aimed at mitigating these issues, organized across the four phases of LLM development and usage: data collecting and pre-training, fine-tuning and alignment, prompting and reasoning, and post-processing and auditing. We elaborate on the recent advances for enhancing the performance of LLMs in terms of privacy protection, hallucination reduction, value alignment, toxicity elimination, and jailbreak defenses. In contrast to previous surveys that focus on a single dimension of responsible LLMs, this survey presents a unified framework that encompasses these diverse dimensions, providing a comprehensive view of enhancing LLMs to better serve real-world applications.

📄 PDF Abstract BibTeX arXiv:2501.09431

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationSurvey

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

A Survey on Data Security in Large Language Models

2025-08-04 · Kang Chen, Xiuze Zhou, Yuanguo Lin, Jinhe Su 외 arxiv

Large Language Models (LLMs), now a foundation in advancing natural language processing, power applications such as text generation, machine translation, and conversational systems. Despite their transformative potential…

Machine TranslationData AugmentationText Generation

A Survey of Attacks on Large Language Models

2025-05-18 · Wenrui Xu, Keshab K. Parhi

Large language models (LLMs) and LLM-based agents have been widely deployed in a wide range of applications in the real world, including healthcare diagnostics, financial analysis, customer support, robotics, and autonom…

Autonomous DrivingFinancial AnalysisSurvey

Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models

2025-09-02 · Ranjie Duan, Jiexi Liu, Xiaojun Jia, Shiji Zhao 외 arxiv

Large language models (LLMs) typically deploy safety mechanisms to prevent harmful content generation. Most current approaches focus narrowly on risks posed by malicious actors, often framing risks as adversarial events …

Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements

2023-02-18 · Jiawen Deng, Jiale Cheng, Hao Sun, Zhexin Zhang 외

As generative large model capabilities advance, safety concerns become more pronounced in their outputs. To ensure the sustainable growth of the AI ecosystem, it's imperative to undertake a holistic evaluation and refine…

Adversarial AttackEthicsSurvey

Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak

2023-12-07 · Yanrui Du, Sendong Zhao, Ming Ma, Yuhan Chen 외

Extensive work has been devoted to improving the safety mechanism of Large Language Models (LLMs). However, LLMs still tend to generate harmful responses when faced with malicious instructions, a phenomenon referred to a…