paper-with-me

홈 › Papers

Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution

2026-03-05 · Qiao Jin, Yin Fang, Lauren He, Yifan Yang, Guangzhi Xiong, Zhizheng Wang, Nicholas Wan, Joey Chan, Donald C. Comeau, Robert Leaman, Charalampos S. Floudas, Aidong Zhang, Michael F. Chiang, Yifan Peng, Zhiyong Lu arxiv

Assessing whether an article supports an assertion is essential for hallucination detection and claim verification. While large language models (LLMs) have the potential to automate this task, achieving strong performance requires frontier models such as GPT-5 that are prohibitively expensive to deploy at scale. To efficiently perform biomedical evidence attribution, we present Med-V1, a family of small language models with only three billion parameters. Trained on high-quality synthetic data newly developed in this study, Med-V1 substantially outperforms (+27.0% to +71.3%) its base models on five biomedical benchmarks unified into a verification format. Despite its smaller size, Med-V1 performs comparably to frontier LLMs such as GPT-5, along with high-quality explanations for its predictions. We use Med-V1 to conduct a first-of-its-kind use case study that quantifies hallucinations in LLM-generated answers under different citation instructions. Results show that the format instruction strongly affects citation validity and hallucination, with GPT-5 generating more claims but exhibiting hallucination rates similar to GPT-4o. Additionally, we present a second use case showing that Med-V1 can automatically identify high-stakes evidence misattributions in clinical practice guidelines, revealing potentially negative public health impacts that are otherwise challenging to identify at scale. Overall, Med-V1 provides an efficient and accurate lightweight alternative to frontier LLMs for practical and real-world applications in biomedical evidence attribution and verification tasks. Med-V1 is available at https://github.com/ncbi-nlp/Med-V1.

📄 PDF Abstract BibTeX arXiv:2603.05308

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluation of ChatGPT on Biomedical Tasks: A Zero-Shot Comparison with Fine-Tuned Generative Transformers

2023-06-07 · Israt Jahan, Md Tahmid Rahman Laskar, Chun Peng, Jimmy Huang

ChatGPT is a large language model developed by OpenAI. Despite its impressive performance across various tasks, no prior work has investigated its capability in the biomedical domain yet. To this end, this paper aims to …

Document ClassificationLanguage ModelingLanguage ModellingLarge Language Model+2

Small LLMs for Biomedical Claim Verification: Cost-Effective Fine-Tuning, Structural Dataset Shortcuts, and Cross-Domain Generalization

2026-06-11 · Gaurav Kumar arxiv

Large Language Models such as GPT-4o and GPT-5 achieve strong zero-shot performance on biomedical claim verification, but cost and opacity limit scalable use. We fine-tune three small LLMs: Phi-3-mini (3.8B), Qwen2.5-3B,…

Domain Generalization

BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing

2022-06-30 · Jason Alan Fries, Leon Weber, Natasha Seelam, Gabriel Altay 외

Training and evaluating language models increasingly requires the construction of meta-datasets --diverse collections of curated data with clear provenance. Natural language prompting has recently lead to improved zero-s…

DiversityLanguage Model EvaluationLanguage ModelingLanguage Modelling+4

Utilizing Large Language Models for Zero-Shot Medical Ontology Extension from Clinical Notes

2025-11-20 · Guanchen Wu, Yuzhang Xie, Huanwei Wu, Zhe He 외 arxiv

Integrating novel medical concepts and relationships into existing ontologies can significantly enhance their coverage and utility for both biomedical research and clinical applications. Clinical notes, as unstructured d…

AutoGraphex: Zero-shot Biomedical Definition Generation with Automatic Prompting

2021-12-17 · ACL ARR December 2022 12 · Anonymous

Describing terminologies with definition texts is an important step towards understanding the scientific literature, especially for domains with limited labeled terminologies. Previous works have sought to design supervi…

DescriptiveLanguage ModelingLanguage ModellingText Generation+1