paper-with-me

Papers

Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)

2026-01-16 · Anjanava Biswas, Wrick Talukdar arxiv

The AI era has ushered in Large Language Models (LLM) to the technological forefront, which has been much of the talk in 2023, and is likely to remain as such for many years to come. LLMs are the AI models that are the power house behind generative AI applications such as ChatGPT. These AI models, fueled by vast amounts of data and computational prowess, have unlocked remarkable capabilities, from human-like text generation to assisting with natural language understanding (NLU) tasks. They have quickly become the foundation upon which countless applications and software services are being built, or at least being augmented with. However, as with any groundbreaking innovations, the rise of LLMs brings forth critical safety, privacy, and ethical concerns. These models are found to have a propensity to leak private information, produce false information, and can be coerced into generating content that can be used for nefarious purposes by bad actors, or even by regular users unknowingly. Implementing safeguards and guardrailing techniques is imperative for applications to ensure that the content generated by LLMs are safe, secure, and ethical. Thus, frameworks to deploy mechanisms that prevent misuse of these models via application implementations is imperative. In this study, wepropose a Flexible Adaptive Sequencing mechanism with trust and safety modules, that can be used to implement safety guardrails for the development and deployment of LLMs.

📄 PDF Abstract BibTeX arXiv:2601.14298

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language UnderstandingText Generation

Similar Papers 제목 키워드 기반

`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts

2025-07-01 · Annika M Schoene, Cansu Canca arxiv

Recent advances in large language models (LLMs) have led to increasingly sophisticated safety protocols and features designed to prevent harmful, unethical, or unauthorized outputs. However, these guardrails remain susce…

Current state of LLM Risks and AI Guardrails

2024-06-16 · Suriya Ganesh Ayyamperumal, Limin Ge

Large language models (LLMs) have become increasingly sophisticated, leading to widespread deployment in sensitive applications where safety and reliability are paramount. However, LLMs have inherent risks accompanying t…

FairnessRAGRetrieval-augmented Generation

From Framework to Reliable Practice: End-User Perspectives on Social Robots in Public Spaces

2025-11-13 · Samson Oruma, Ricardo Colomo-Palacios, Vasileios Gkioulos arxiv

As social robots increasingly enter public environments, their acceptance depends not only on technical reliability but also on ethical integrity, accessibility, and user trust. This paper reports on a pilot deployment o…

Trust-Oriented Adaptive Guardrails for Large Language Models

2024-08-16 · Jinwei Hu, Yi Dong, Xiaowei Huang

Guardrail, an emerging mechanism designed to ensure that large language models (LLMs) align with human values by moderating harmful or toxic responses, requires a sociotechnical approach in their design. This paper addre…

In-Context LearningRetrieval-augmented Generation

AI Ethics by Design: Implementing Customizable Guardrails for Responsible AI Development

2024-11-05 · Kristina Šekrst, Jeremy McHugh, Jonathan Rodriguez Cefalu

This paper explores the development of an ethical guardrail framework for AI systems, emphasizing the importance of customizable guardrails that align with diverse user values and underlying ethics. We address the challe…

Ethics