paper-with-me

Papers

RAmBLA: A Framework for Evaluating the Reliability of LLMs as Assistants in the Biomedical Domain

2024-03-21 · William James Bolton, Rafael Poyiadzi, Edward R. Morrell, Gabriela van Bergen Gonzalez Bueno, Lea Goetz

Large Language Models (LLMs) increasingly support applications in a wide range of domains, some with potential high societal impact such as biomedicine, yet their reliability in realistic use cases is under-researched. In this work we introduce the Reliability AssesMent for Biomedical LLM Assistants (RAmBLA) framework and evaluate whether four state-of-the-art foundation LLMs can serve as reliable assistants in the biomedical domain. We identify prompt robustness, high recall, and a lack of hallucinations as necessary criteria for this use case. We design shortform tasks and tasks requiring LLM freeform responses mimicking real-world user interactions. We evaluate LLM performance using semantic similarity with a ground truth response, through an evaluator LLM.

📄 PDF Abstract BibTeX arXiv:2403.14578

Code (1)

gsk-ai/rambla 공식 구현

Tasks

Semantic SimilaritySemantic Textual Similarity

Similar Papers 제목 키워드 기반

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory

2026-06-02 · Noujoud Nader, Ibrahem Aljabea, Patrick Diehl, Deepti Gupta arxiv

Large language models (LLMs) are increasingly used as self-study assistants in technical disciplines, yet their reliability as mathematical reasoning assistants remains poorly understood. We introduce GTBench, a curricul…

Mathematical Reasoning

Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators

2024-09-21 · Prasoon Bajpai, Niladri Chatterjee, Subhabrata Dutta, Tanmoy Chakraborty

Large Language Models (LLMs) and AI assistants driven by these models are experiencing exponential growth in usage among both expert and amateur users. In this work, we focus on evaluating the reliability of current LLMs…

Benchmarking

Depth and Autonomy: A Framework for Evaluating LLM Applications in Social Science Research

2025-10-29 · Ali Sanaei, Ali Rajabzadeh arxiv

Large language models (LLMs) are increasingly utilized by researchers across a wide range of domains, and qualitative social science is no exception; however, this adoption faces persistent challenges, including interpre…

Phosphatidylserine transport in cell life and death

2023-11-09 · Alenka {Č}opi{č}, Thibaud Dieudonné, Guillaume Lenoir

Phosphatidylserine (PS) is a negatively-charged glycerophospholipid found mainly in the plasma membrane (PM) and in the late secretory/endocytic compartments, where it regulates cellular activity and can mediate apoptosi…

Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards

2026-04-25 · Taha Hammadia, Lucas Rea, Ahmad Mohammad Saber, Amr Youssef 외 arxiv

The deployment of Large Language Models (LLMs) as assistants in electric grid operations promises to streamline compliance and decision-making but exposes new vulnerabilities to prompt-based adversarial attacks. This pap…