paper-with-me

홈 › Papers

Alignment Plausibility: A New Standard for Assuring AI in Healthcare

2026-07-08 · Gwydion Williams, Sara Zannone, Bilal A Mateen arxiv

Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers' safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.g., dependency, boundary erosion, the amplification of distorted beliefs) receive less attention. We contend that making LLMs structurally safe requires alignment organised at three levels that mirror how society assures the safety of human clinical practice: 1) explicit value specification grounded in the codified normative commitments of clinical practice; 2) training that embeds those values in the model; and 3) oversight that detects drift and longer-term harm during deployment, much as clinical supervision does for human practice. Organising alignment in this way yields a construct we call alignment plausibility - a structured demonstration that a system's values, training regime, and oversight mechanisms are together consistent with safe and positive outcomes. We propose alignment plausibility as a regulatory construct (by drawing analogy to the established construct of biological plausibility) for AI in health: a principled way to argue for, or against, trust that systems are aligned to positive health outcomes, will cause no harm even where capable of doing so, and will ultimately lead to patient benefit.

📄 PDF Abstract BibTeX arXiv:2607.07766

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Foundational Challenges in Assuring Alignment and Safety of Large Language Models

2024-04-15 · Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka 외

This work identifies 18 foundational challenges in assuring the alignment and safety of large language models (LLMs). These challenges are organized into three different categories: scientific understanding of LLMs, deve…

The Role of Explainability in Assuring Safety of Machine Learning in Healthcare

2021-09-01 · Yan Jia, John McDermid, Tom Lawton, Ibrahim Habli

Established approaches to assuring safety-critical systems and software are difficult to apply to systems employing ML where there is no clear, pre-defined specification against which to assess validity. This problem is …

BIG-bench Machine LearningExplainable Artificial Intelligence (XAI)

A Framework for Human Evaluation of Large Language Models in Healthcare Derived from Literature Review

2024-05-04 · Thomas Yu CHow Tam, Sonish Sivarajkumar, Sumit Kapoor, Alisa V Stolyar 외

With generative artificial intelligence (AI), particularly large language models (LLMs), continuing to make inroads in healthcare, it is critical to supplement traditional automated evaluations with human evaluations. Un…

Biological Plausibility and Representational Alignment of Feedback Alignment in Convolutional Networks

2026-05-08 · Jake Lance, Larry Kieu arxiv

The feedback alignment (FA) algorithm offers a biologically plausible alternative to backpropagation (BP) for training neural networks yet notably fails to scale to convolutional architectures. Modifications have been pr…

Review of the AMLAS Methodology for Application in Healthcare

2022-09-01 · Shakir Laher, Carla Brackstone, Sara Reis, An Nguyen 외

In recent years, the number of machine learning (ML) technologies gaining regulatory approval for healthcare has increased significantly allowing them to be placed on the market. However, the regulatory frameworks applie…