paper-with-me

홈 › Papers

M-QALM: A Benchmark to Assess Clinical Reading Comprehension and Knowledge Recall in Large Language Models via Question Answering

2024-06-06 · Anand Subramanian, Viktor Schlegel, Abhinav Ramesh Kashyap, Thanh-Tung Nguyen, Vijay Prakash Dwivedi, Stefan Winkler

There is vivid research on adapting Large Language Models (LLMs) to perform a variety of tasks in high-stakes domains such as healthcare. Despite their popularity, there is a lack of understanding of the extent and contributing factors that allow LLMs to recall relevant knowledge and combine it with presented information in the clinical and biomedical domain: a fundamental pre-requisite for success on down-stream tasks. Addressing this gap, we use Multiple Choice and Abstractive Question Answering to conduct a large-scale empirical study on 22 datasets in three generalist and three specialist biomedical sub-domains. Our multifaceted analysis of the performance of 15 LLMs, further broken down by sub-domain, source of knowledge and model architecture, uncovers success factors such as instruction tuning that lead to improved recall and comprehension. We further show that while recently proposed domain-adapted models may lack adequate knowledge, directly fine-tuning on our collected medical knowledge datasets shows encouraging results, even generalising to unseen specialist sub-domains. We complement the quantitative results with a skill-oriented manual error analysis, which reveals a significant gap between the models' capabilities to simply recall necessary knowledge and to integrate it with the presented context. To foster research and collaboration in this field we share M-QALM, our resources, standardised methodology, and evaluation results, with the research community to facilitate further advancements in clinical knowledge representation learning within language models.

📄 PDF Abstract BibTeX arXiv:2406.03699

Code (1)

anand-subu/m-qalm 공식 구현

Tasks

abstractive question answeringClinical KnowledgeMultiple-choiceQuestion AnsweringReading ComprehensionRepresentation Learning

Similar Papers 제목 키워드 기반

A Survey of Machine Narrative Reading Comprehension Assessments

2022-04-30 · Yisi Sang, Xiangyang Mou, Jing Li, Jeffrey Stanton 외

As the body of research on machine narrative comprehension grows, there is a critical need for consideration of performance assessment strategies as well as the depth and scope of different benchmark tasks. Based on narr…

Reading ComprehensionSurvey

RoBIn: A Transformer-Based Model For Risk Of Bias Inference With Machine Reading Comprehension

2024-10-28 · Abel Corrêa Dias, Viviane Pereira Moreira, João Luiz Dihl Comba

Objective: Scientific publications play a crucial role in uncovering insights, testing novel drugs, and shaping healthcare policies. Accessing the quality of publications requires evaluating their Risk of Bias (RoB), a p…

Binary ClassificationMachine Reading ComprehensionReading Comprehension

MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark

2025-03-10 · Shengkun Ma, Hao Peng, Lei Hou, Juanzi Li

Machine Reading Comprehension (MRC) is an essential task in evaluating natural language understanding. Existing MRC datasets primarily assess specific aspects of reading comprehension (RC), lacking a comprehensive MRC be…

Machine Reading ComprehensionNatural Language UnderstandingReading Comprehension

Automatic learner summary assessment for reading comprehension

2019-06-18 · NAACL 2019 6 · Menglin Xia, Ekaterina Kochmar, Ted Briscoe

Automating the assessment of learner summaries provides a useful tool for assessing learner reading comprehension. We present a summarization task for evaluating non-native reading comprehension and propose three novel a…

Reading Comprehension

Analyzing Multiple-Choice Reading and Listening Comprehension Tests

2023-07-03 · Vatsal Raina, Adian Liusie, Mark Gales

Multiple-choice reading and listening comprehension tests are an important part of language assessment. Content creators for standard educational tests need to carefully curate questions that assess the comprehension abi…

Multiple-choiceReading ComprehensionWorld Knowledge