paper-with-me

홈 › Papers

CLIMB: A Benchmark of Clinical Bias in Large Language Models

2024-07-07 · Yubo Zhang, Shudi Hou, Mingyu Derek Ma, Wei Wang, Muhao Chen, Jieyu Zhao

Large language models (LLMs) are increasingly applied to clinical decision-making. However, their potential to exhibit bias poses significant risks to clinical equity. Currently, there is a lack of benchmarks that systematically evaluate such clinical bias in LLMs. While in downstream tasks, some biases of LLMs can be avoided such as by instructing the model to answer "I'm not sure...", the internal bias hidden within the model still lacks deep studies. We introduce CLIMB (shorthand for A Benchmark of Clinical Bias in Large Language Models), a pioneering comprehensive benchmark to evaluate both intrinsic (within LLMs) and extrinsic (on downstream tasks) bias in LLMs for clinical decision tasks. Notably, for intrinsic bias, we introduce a novel metric, AssocMAD, to assess the disparities of LLMs across multiple demographic groups. Additionally, we leverage counterfactual intervention to evaluate extrinsic bias in a task of clinical diagnosis prediction. Our experiments across popular and medically adapted LLMs, particularly from the Mistral and LLaMA families, unveil prevalent behaviors with both intrinsic and extrinsic bias. This work underscores the critical need to mitigate clinical bias and sets a new standard for future evaluations of LLMs' clinical bias.

📄 PDF Abstract BibTeX arXiv:2407.05250

Code (1)

uscnlp-lime/climb 공식 구현

Tasks

counterfactualDecision Making

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models

2025-03-09 · Wei Dai, Peilin Chen, Malinda Lu, Daniel Li 외

Recent advances in clinical AI have enabled remarkable progress across many clinical domains. However, existing benchmarks and models are primarily limited to a small set of modalities and tasks, which hinders the develo…

CliMB: An AI-enabled Partner for Clinical Predictive Modeling

2024-09-30 · Evgeny Saveliev, Tim Schubert, Thomas Pouplin, Vasilis Kosmoliaptsis 외

Despite its significant promise and continuous technical advances, real-world applications of artificial intelligence (AI) remain limited. We attribute this to the "domain expert-AI-conundrum": while domain experts, such…

AttributeAutoML

CLiMB: A Continual Learning Benchmark for Vision-and-Language Tasks

2022-06-18 · Tejas Srinivasan, Ting-Yun Chang, Leticia Leonor Pinto Alva, Georgios Chochlakis 외

Current state-of-the-art vision-and-language models are evaluated on tasks either individually or in a multi-task setting, overlooking the challenges of continually learning (CL) tasks as they arrive. Existing CL benchma…

Continual LearningTransfer Learning

CLIMB grammars: three projects using metagrammar engineering

2012-05-01 · LREC 2012 5 · Antske Fokkens, Tania Avgustinova, Yi Zhang

This paper introduces the CLIMB (Comparative Libraries of Implementations with Matrix Basis) methodology and grammars. The basic idea behind CLIMB is to use code generation as a general methodology for grammar developmen…

Code Generation

Board-to-Board: Evaluating Moonboard Grade Prediction Generalization

2023-11-21 · Daniel Petashvili, Matthew Rodda

Bouldering is a sport where athletes aim to climb up an obstacle using a set of defined holds called a route. Typically routes are assigned a grade to inform climbers of its difficulty and allow them to more easily track…

Prediction