Yesil o1 Pro: Evidence-Based AI Model for Health and Benchmarking in Clinical Decision Support
Background: Integrating evidence-based approaches in healthcare and artificial intelligence (AI) is crucial for enhancing clinical decision-making and patient safety. Yesil o1 Pro is a specialized large language model (LLM) designed to transform medical knowledge synthesis by leveraging a comprehensive, curated database of scientific literature, clinical guidelines, and medical textbooks. Objective: The system's innovative "AI Hospital" framework employs domain-specific expert agents coordinated by a central Master Agent, enabling tailored and precise medical responses across multiple disciplines. Methods: The model's advanced methodology incorporates sophisticated techniques including GraphRAG-based retrieval, extensive fine-tuning with 1.5 million question-answer pairs, and Chain of Thought (CoT) reasoning. Its robust training dataset comprises 100.5M words from high-impact journals, 96.8M words from core medical texts, and 74.4M words from international standards, ensuring a comprehensive and authoritative knowledge base. Results: Benchmark evaluations demonstrate Yesil o1 Pro's exceptional performance, achieving an overall accuracy of 89.1% and surpassing leading models like GPT-4o (83.9%) and Claude 3.5 Sonnet (83.0%). Domain-specific accuracies are particularly impressive, with 96.1% in Mental Health, 94.6% in Epidemiology, and 94.6% in Dentistry, highlighting the model's proficiency in handling complex, reasoning-intensive medical queries. Conclusions: The model shows promising applications in clinical decision support, interdisciplinary collaboration, professional education, and medical research. While challenges remain in real-world integration and maintaining alignment with evolving medical knowledge, Yesil o1 Pro represents a significant advancement in AI-driven healthcare support. Future research will focus on validating the model's utility in clinical environments and developing strategies for seamless healthcare system integration.
Code (0)
등록된 구현이 없습니다.
Tasks
BenchmarkingEpidemiologyLarge Language ModelMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Accurate, fully-automated NMR spectral profiling for metabolomics
Many diseases cause significant changes to the concentrations of small molecules (aka metabolites) that appear in a person's biofluids, which means such diseases can often be readily detected from a person's "metabolic p…
CPUMedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization
Real-world clinical decision support requires reasoning over heterogeneous and longitudinal patient information rather than answering isolated medical questions. However, current medical large language models and retriev…
Question AnsweringFrom Questions to Clinical Recommendations: Large Language Models Driving Evidence-Based Clinical Decision Making
Clinical evidence, derived from rigorous research and data analysis, provides healthcare professionals with reliable scientific foundations for informed decision-making. Integrating clinical evidence into real-time pract…
Decision MakingDeepER-Med: Advancing Deep Evidence-Based Research in Medicine Through Agentic AI
Trustworthiness and transparency are essential for the clinical adoption of artificial intelligence (AI) in healthcare and biomedical research. Recent deep research systems aim to accelerate evidence-grounded scientific …
Information RetrievalMamaBench: Benchmarking LLM Robustness in Maternal and Child Health Diagnosis through Counterfactual Clinical Perturbation
Large language models achieve strong scores on medical benchmarks, yet these benchmarks evaluate each question in isolation, providing no measure of whether a system can distinguish clinically similar presentations requi…