paper-with-me

홈 › Papers

Helpful or Harmful? Exploring the Efficacy of Large Language Models for Online Grooming Prevention

2024-03-14 · Ellie Prosser, Matthew Edwards

Powerful generative Large Language Models (LLMs) are becoming popular tools amongst the general public as question-answering systems, and are being utilised by vulnerable groups such as children. With children increasingly interacting with these tools, it is imperative for researchers to scrutinise the safety of LLMs, especially for applications that could lead to serious outcomes, such as online child safety queries. In this paper, the efficacy of LLMs for online grooming prevention is explored both for identifying and avoiding grooming through advice generation, and the impact of prompt design on model performance is investigated by varying the provided context and prompt specificity. In results reflecting over 6,000 LLM interactions, we find that no models were clearly appropriate for online grooming prevention, with an observed lack of consistency in behaviours, and potential for harmful answer generation, especially from open-source models. We outline where and how models fall short, providing suggestions for improvement, and identify prompt designs that heavily altered model performance in troubling ways, with findings that can be used to inform best practice usage guides.

📄 PDF Abstract BibTeX arXiv:2403.09795

Code (0)

등록된 구현이 없습니다.

Tasks

Answer GenerationQuestion AnsweringSpecificity

Similar Papers 제목 키워드 기반

Fluid Transformers and Creative Analogies: Exploring Large Language Models' Capacity for Augmenting Cross-Domain Analogical Creativity

2023-02-27 · Zijian Ding, Arvind Srinivasan, Stephen MacNeil, Joel Chan

Cross-domain analogical reasoning is a core creative ability that can be challenging for humans. Recent work has shown some proofs-of concept of Large language Models' (LLMs) ability to generate cross-domain analogies. H…

How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation

2025-02-20 · Zhuohang Long, Siyuan Wang, Shujun Liu, Yuhang Lai 외

Jailbreak attacks, where harmful prompts bypass generative models' built-in safety, raise serious concerns about model vulnerability. While many defense methods have been proposed, the trade-offs between safety and helpf…

Binary Classification

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions

2025-02-08 · Jingxin Xu, Guoshun Nan, Sheng Guan, Sicong Leng 외

Recent AI agents, such as ChatGPT and LLaMA, primarily rely on instruction tuning and reinforcement learning to calibrate the output of large language models (LLMs) with human intentions, ensuring the outputs are harmles…

Safety Alignment

Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting

2025-10-02 · Andrea Wynn, Metod Jazbec, Charith Peris, Rinat Khaziev 외 arxiv

Large language models (LLMs) can be influenced by harmful or irrelevant context, which can significantly harm model performance on downstream tasks. This motivates principled designs in which LLM systems include built-in…

Computational EfficiencyQuestion Answering

Navigating the OverKill in Large Language Models

2024-01-31 · Chenyu Shi, Xiao Wang, Qiming Ge, Songyang Gao 외

Large language models are meticulously aligned to be both helpful and harmless. However, recent research points to a potential overkill which means models may refuse to answer benign queries. In this paper, we investigat…