paper-with-me

Papers

Context Matters: Comparison of commercial large language tools in veterinary medicine

2025-09-22 · Tyler J Poore, Christopher J Pinard, Aleena Shabbir, Andrew Lagree, Andre Telfer, Kuan-Chuen Wu arxiv

Large language models (LLMs) are increasingly used in clinical settings, yet their performance in veterinary medicine remains underexplored. We evaluated three commercially available veterinary-focused LLM summarization tools (Product 1 [Hachiko] and Products 2 and 3) on a standardized dataset of veterinary oncology records. Using a rubric-guided LLM-as-a-judge framework, summaries were scored across five domains: Factual Accuracy, Completeness, Chronological Order, Clinical Relevance, and Organization. Product 1 achieved the highest overall performance, with a median average score of 4.61 (IQR: 0.73), compared to 2.55 (IQR: 0.78) for Product 2 and 2.45 (IQR: 0.92) for Product 3. It also received perfect median scores in Factual Accuracy and Chronological Order. To assess the internal consistency of the grading framework itself, we repeated the evaluation across three independent runs. The LLM grader demonstrated high reproducibility, with Average Score standard deviations of 0.015 (Product 1), 0.088 (Product 2), and 0.034 (Product 3). These findings highlight the importance of veterinary-specific commercial LLM tools and demonstrate that LLM-as-a-judge evaluation is a scalable and reproducible method for assessing clinical NLP summarization in veterinary medicine.

📄 PDF Abstract BibTeX arXiv:2510.01224

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs

2025-08-15 · Mikhail Seleznyov, Mikhail Chaichuk, Gleb Ershov, Alexander Panchenko 외 arxiv

Large Language Models (LLMs) are highly sensitive to subtle, non-semantic variations in prompt phrasing and formatting. In this work, we present the first systematic evaluation of 5 methods for improving prompt robustnes…

Redact or Keep? A Fully Local AI Cascade for Educational Dialogue De-Identification

2026-06-16 · Haocheng Zhang, Zhuqian Zhou, Kirk Vanacore, Bakhtawar Ahtisham 외 arxiv

Educational dialogue is a valuable but sensitive resource for research: the same transcripts that capture authentic learning often capture personally identifiable information (PII) entangled with curricular content, wher…

Keyword Augmented Retrieval: Novel framework for Information Retrieval integrated with speech interface

2023-10-06 · Anupam Purwar, Rahul Sundar

Retrieving answers in a quick and low cost manner without hallucinations from a combination of structured and unstructured data using Language models is a major hurdle. This is what prevents employment of Language models…

ChatbotInformation RetrievalLanguage ModellingLarge Language Model+1

Türkçe Dil Modellerinin Performans Karşılaştırması Performance Comparison of Turkish Language Models

2024-04-25 · Eren Dogan, M. Egemen Uzun, Atahan Uz, H. Emre Seyrek 외

The developments that language models have provided in fulfilling almost all kinds of tasks have attracted the attention of not only researchers but also the society and have enabled them to become products. There are co…

In-Context LearningQuestion Answering

How Good are Commercial Large Language Models on African Languages?

2023-05-11 · Jessica Ojo, Kelechi Ogueji

Recent advancements in Natural Language Processing (NLP) has led to the proliferation of large pretrained language models. These models have been shown to yield good performance, using in-context learning, even on unseen…

In-Context LearningLanguage ModelingLanguage ModellingMachine Translation+3