paper-with-me

홈 › Papers

AI Safety in Generative AI Large Language Models: A Survey

2024-07-06 · Jaymari Chua, Yun Li, Shiyi Yang, Chen Wang, Lina Yao

Large Language Model (LLMs) such as ChatGPT that exhibit generative AI capabilities are facing accelerated adoption and innovation. The increased presence of Generative AI (GAI) inevitably raises concerns about the risks and safety associated with these models. This article provides an up-to-date survey of recent trends in AI safety research of GAI-LLMs from a computer scientist's perspective: specific and technical. In this survey, we explore the background and motivation for the identified harms and risks in the context of LLMs being generative language models; our survey differentiates by emphasising the need for unified theories of the distinct safety challenges in the research development and applications of LLMs. We start our discussion with a concise introduction to the workings of LLMs, supported by relevant literature. Then we discuss earlier research that has pointed out the fundamental constraints of generative models, or lack of understanding thereof (e.g., performance and safety trade-offs as LLMs scale in number of parameters). We provide a sufficient coverage of LLM alignment -- delving into various approaches, contending methods and present challenges associated with aligning LLMs with human preferences. By highlighting the gaps in the literature and possible implementation oversights, our aim is to create a comprehensive analysis that provides insights for addressing AI safety in LLMs and encourages the development of aligned and secure models. We conclude our survey by discussing future directions of LLMs for AI safety, offering insights into ongoing research in this critical area.

📄 PDF Abstract BibTeX arXiv:2407.18369

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModellingLarge Language ModelSurvey

Similar Papers 제목 키워드 기반

Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements

2023-02-18 · Jiawen Deng, Jiale Cheng, Hao Sun, Zhexin Zhang 외

As generative large model capabilities advance, safety concerns become more pronounced in their outputs. To ensure the sustainable growth of the AI ecosystem, it's imperative to undertake a holistic evaluation and refine…

Adversarial AttackEthicsSurvey

Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)

2024-07-10 · Krishnaram Kenthapadi, Mehrnoosh Sameki, Ankur Taly

With the ongoing rapid adoption of Artificial Intelligence (AI)-based systems in high-stakes domains, ensuring the trustworthiness, safety, and observability of these systems has become crucial. It is essential to evalua…

Survey

Evaluating the Bias in LLMs for Surveying Opinion and Decision Making in Healthcare

2025-04-11 · Yonchanok Khaokaew, Flora D. Salim, Andreas Züfle, Hao Xue 외

Generative agents have been increasingly used to simulate human behaviour in silico, driven by large language models (LLMs). These simulacra serve as sandboxes for studying human behaviour without compromising privacy or…

Decision MakingPrompt EngineeringSurvey

LLMs Meet Multimodal Generation and Editing: A Survey

2024-05-29 · Yingqing He, Zhaoyang Liu, Jingye Chen, Zeyue Tian 외

With the recent advancement in large language models (LLMs), there is a growing interest in combining LLMs with multimodal learning. Previous surveys of multimodal large language models (MLLMs) mainly focus on multimodal…

multimodal generationSurvey

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

2025-06-05 · Seongmin Lee, Aeree Cho, Grace C. Kim, Shengyun Peng 외

As large language models (LLMs) see wider real-world use, understanding and mitigating their unsafe behaviors is critical. Interpretation techniques can reveal causes of unsafe outputs and guide safety, but such connecti…

NavigateSurvey