paper-with-me

홈 › Papers

PeruMedQA: Benchmarking Large Language Models (LLMs) on Peruvian Medical Exams -- Dataset Construction and Evaluation

2025-09-15 · Rodrigo M. Carrillo-Larco, Jesus Lovón Melgarejo, Manuel Castillo-Cara, Gusseppe Bravo-Rocca arxiv

BACKGROUND: Medical large language models (LLMs) have demonstrated remarkable performance in answering medical examinations. However, the extent to which this high performance is transferable to medical questions in Spanish and from a Latin American country remains unexplored. This knowledge is crucial as LLM-based medical applications gain traction in Latin America. AIMS: To build a dataset of questions medical examinations taken by Peruvian physicians pursuing specialty training; to fine-tune a LLM on this dataset; to evaluate and compare the performance in terms of accuracy between vanilla LLMs and the fine-tuned LLM. METHODS: We curated PeruMedQA, a multiple-choice question-answering (MCQA) dataset containing 8,380 questions spanning 12 specialties (2018-2025). We selected ten medical LLMs, including medgemma-4b-it and medgemma-27b-text-it, and developed zero-shot task specific prompts to answer the questions. We employed parameter-efficient fine tuning (PEFT) and low-rand adaptation (LoRA) to fine-tune medgemma-4b-it utilizing all questions except those from 2025 (test set). RESULTS: Medgemma-27b showed the highest accuracy across all specialities, achieving the highest score of 89.29% in Psychiatry; yet, in two specialties, OctoMed-7B exhibited slight superiority: Neurosurgery with 77.27% and 77.38, respectively; and Radiology with 76.13% and 77.39%, respectively. Across specialties, most LLMs with <10 billion parameters exhibited <50% of correct answers. The fine-tuned version of medgemma-4b-it emerged victorious against all LLMs with <10 billion parameters and rivaled a LLM with 70 billion parameters across various examinations. CONCLUSIONS: For medical AI applications and research that require knowledge bases from Spanish-speaking countries and those exhibiting similar epidemiological profile to Peru's, interested parties should utilize medgemma-27b-text-it.

📄 PDF Abstract BibTeX arXiv:2509.11517

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WordNet-Shp: Towards the Building of a Lexical Database for a Peruvian Minority Language

2018-05-01 · LREC 2018 5 · Diego Magui{\~n}o-Valencia, Arturo Oncevay-Marcos, Marco A. Sobrevilla Cabezudo
Machine TranslationWord Sense Disambiguation

Text Mining over Curriculum Vitae of Peruvian Professionals using Official Scientific Site DINA

2020-09-07 · Josimar Edinson Chire Saire, Honorio Apaza Alanoca

During the last decade, Peruvian government started to invest and promote Science and Technology through Concytec(National Council of Science and Technology). Many programs are oriented to support research projects, expe…

Enhancing Cross Domain SAR Oil Spill Segmentation via Morphological Region Perturbation and Synthetic Label-to-SAR Generation

2025-12-02 · Andre Juarez, Luis Salsavilca, Frida Coaquira, Celso Gonzales arxiv

Deep learning models for SAR oil spill segmentation often fail to generalize across regions due to differences in sea-state, backscatter statistics, and slick morphology, a limitation that is particularly severe along th…

Lee-Carter method for forecasting mortality for Peruvian Population

2018-11-23

In this article, we have modeled mortality rates of Peruvian female and male populations during the period of 1950-2017 using the Lee-Carter (LC) model. The stochastic mortality model was introduced by Lee and Carter (19…

PeruSIL: A Framework to Build a Continuous Peruvian Sign Language Interpretation Dataset

2022-06-01 · SignLang (LREC) 2022 6 · Gissella Bejarano, Joe Huamani-Malca, Francisco Cerna-Herrera, Fernando Alva-Manchego 외

Video-based datasets for Continuous Sign Language are scarce due to the challenging task of recording videos from native signers and the reduced number of people who can annotate sign language. COVID-19 has evidenced the…

Sign Language Recognition