paper-with-me

Papers

Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks

2025-08-13 · Nouar AlDahoul, Yasir Zaki arxiv

Recent progress in large language models (LLMs) has showcased impressive proficiency in numerous Arabic natural language processing (NLP) applications. Nevertheless, their effectiveness in Arabic medical NLP domains has received limited investigation. This research examines the degree to which state-of-the-art LLMs demonstrate and articulate healthcare knowledge in Arabic, assessing their capabilities across a varied array of Arabic medical tasks. We benchmark several LLMs using a medical dataset proposed in the Arabic NLP AraHealthQA challenge in MedArabiQ2025 track. Various base LLMs were assessed on their ability to accurately provide correct answers from existing choices in multiple-choice questions (MCQs) and fill-in-the-blank scenarios. Additionally, we evaluated the capacity of LLMs in answering open-ended questions aligned with expert answers. Our results reveal significant variations in correct answer prediction accuracy and low variations in semantic alignment of generated answers, highlighting both the potential and limitations of current LLMs in Arabic clinical contexts. Our analysis shows that for MCQs task, the proposed majority voting solution, leveraging three base models (Gemini Flash 2.5, Gemini Pro 2.5, and GPT o3), outperforms others, achieving up to 77% accuracy and securing first place overall in the Arahealthqa 2025 shared task-track 2 (sub-task 1) challenge. Moreover, for the open-ended questions task, several LLMs were able to demonstrate excellent performance in terms of semantic alignment and achieve a maximum BERTScore of 86.44%.

📄 PDF Abstract BibTeX arXiv:2508.15797

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation

2025-11-27 · Zhen Chen, Yihang Fu, Gabriel Madera, Mauro Giuffre 외 arxiv

Multi-modal large language models (MLLMs) have shown promise in advancing healthcare. However, most existing models remain confined to single-image understanding, which greatly limits their applicability in clinical work…

Medical Diagnosis

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study

2025-07-15 · Che Liu, Jiazhen Pan, Weixiang Shen, Wenjia Bai 외 arxiv

Vision-Language Models (VLMs) trained on web-scale corpora excel at natural image tasks and are increasingly repurposed for healthcare; however, their competence in medical tasks remains underexplored. We present a compr…

U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding

2025-05-23 · Anjie Le, Henan Liu, Yue Wang, Zhenyu Liu 외

Ultrasound is a widely-used imaging modality critical to global healthcare, yet its interpretation remains challenging due to its varying image quality on operators, noises, and anatomical structures. Although large visi…

BenchmarkingSpatial ReasoningText Generation

Benchmarking GPT-5 for biomedical natural language processing

2025-08-28 · Yu Hou, Zaifu Zhan, Min Zeng, Yifan Wu 외 arxiv

Biomedical literature and clinical narratives pose multifaceted challenges for natural language understanding, from precise entity extraction and document synthesis to multi-step diagnostic reasoning. This study extends …

Natural Language UnderstandingDocument ClassificationRelation Extraction

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

2025-07-10 · Shuang Zhou, Wenya Xie, Jiaxi Li, Zaifu Zhan 외 arxiv

As large language models (LLMs) become increasingly integrated into clinical decision-making, ensuring transparent and trustworthy reasoning is essential. However, existing evaluation strategies of LLMs' medical reasonin…