paper-with-me

홈 › Papers

Enhancing Reasoning Skills in Small Persian Medical Language Models Can Outperform Large-Scale Data Training

2025-10-22 · Mehrdad Ghassabi, Sadra Hakim, Hamidreza Baradaran Kashani, Pedram Rostami arxiv

Enhancing reasoning capabilities in small language models is critical for specialized applications such as medical question answering, particularly in underrepresented languages like Persian. In this study, we employ Reinforcement Learning with AI Feedback (RLAIF) and Direct preference optimization (DPO) to improve the reasoning skills of a general-purpose Persian language model. To achieve this, we translated a multiple-choice medical question-answering dataset into Persian and used RLAIF to generate rejected-preferred answer pairs, which are essential for DPO training. By prompting both teacher and student models to produce Chain-of-Thought (CoT) reasoning responses, we compiled a dataset containing correct and incorrect reasoning trajectories. This dataset, comprising 2 million tokens in preferred answers and 2.5 million tokens in rejected ones, was used to train a baseline model, significantly enhancing its medical reasoning capabilities in Persian. Remarkably, the resulting model outperformed its predecessor, gaokerena-V, which was trained on approximately 57 million tokens, despite leveraging a much smaller dataset. These results highlight the efficiency and effectiveness of reasoning-focused training approaches in developing domain-specific language models with limited data availability.

📄 PDF Abstract BibTeX arXiv:2510.20059

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningQuestion Answering

Similar Papers 제목 키워드 기반

Gaokerena: A Small Persian Medical Language Model Family

2026-08-02 · Mehrdad Ghassabi, Hamidreza Baradaran Kashani, Pedram Rostami, Sadra Hakim 외 arxiv

The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused on English, leaving low resource languages like Persian significant…

Reinforcement Learning

PersianMedQA: Language-Centric Evaluation of LLMs in the Persian Medical Domain

2025-05-30 · Mohammad Javad Ranjbar Kalahroodi, Amirhossein Sheikholselami, Sepehr Karimi, Sepideh Ranjbar Kalahroodi 외

Large Language Models (LLMs) have achieved remarkable performance on a wide range of NLP benchmarks, often surpassing human-level accuracy. However, their reliability in high-stakes domains such as medicine, particularly…

Instruction FollowingMultiple-choice

Leveraging Online Data to Enhance Medical Knowledge in a Small Persian Language Model

2025-05-21 · Mehrdad ghassabi, Pedram Rostami, Hamidreza Baradaran Kashani, Amirhossein Poursina 외

The rapid advancement of language models has demonstrated the potential of artificial intelligence in the healthcare industry. However, small language models struggle with specialized domains in low-resource languages li…

Language ModelingLanguage ModellingMedical Question AnsweringPatient QA+2

Benchmarking Large Language Models for Persian: A Preliminary Study Focusing on ChatGPT

2024-04-03 · Amirhossein Abaskohi, Sara Baruni, Mostafa Masoudi, Nesa Abbasi 외

This paper explores the efficacy of large language models (LLMs) for Persian. While ChatGPT and consequent LLMs have shown remarkable performance in English, their efficiency for more low-resource languages remains an op…

BenchmarkingGeneral KnowledgeMath

MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment

2025-08-24 · Omid Ghahroodi, Arshia Hemmat, Marzia Nouri, Seyed Mohammad Hadi Hosseini 외 arxiv

Recent advancements in large vision-language models (VLMs) have primarily focused on English, with limited attention given to other languages. To address this gap, we introduce MEENA (also known as PersianMMMU), the firs…