paper-with-me

홈 › Papers

Performance Evaluation of Open-Source Large Language Models for Assisting Pathology Report Writing in Japanese

2026-03-12 · Masataka Kawai, Singo Sakashita, Shumpei Ishikawa, Shogo Watanabe, Anna Matsuoka, Mikio Sakurai, Yasuto Fujimoto, Yoshiyuki Takahara, Atsushi Ohara, Hirohiko Miyake, Genichiro Ishii arxiv

The performance of large language models (LLMs) for supporting pathology report writing in Japanese remains unexplored. We evaluated seven open-source LLMs from three perspectives: (A) generation and information extraction of pathology diagnosis text following predefined formats, (B) correction of typographical errors in Japanese pathology reports, and (C) subjective evaluation of model-generated explanatory text by pathologists and clinicians. Thinking models and medical-specialized models showed advantages in structured reporting tasks that required reasoning and in typo correction. In contrast, preferences for explanatory outputs varied substantially across raters. Although the utility of LLMs differed by task, our findings suggest that open-source LLMs can be useful for assisting Japanese pathology report writing in limited but clinically relevant scenarios.

📄 PDF Abstract BibTeX arXiv:2603.11597

Code (0)

등록된 구현이 없습니다.

Tasks

Information Extraction

Similar Papers 제목 키워드 기반

Panda LLM: Training Data and Evaluation for Open-Sourced Chinese Instruction-Following Large Language Models

2023-05-04 · Fangkai Jiao, Bosheng Ding, Tianze Luo, Zhanfeng Mo

This project focuses on enhancing open-source large language models through instruction-tuning and providing comprehensive evaluations of their performance. We explore how various training data factors, such as quantity,…

Instruction Following

OpenEthics: A Comprehensive Ethical Evaluation of Open-Source Generative Large Language Models

2025-05-21 · Burak Erinç Çetin, Yıldırım Özen, Elif Naz Demiryılmaz, Kaan Engür 외

Generative large language models present significant potential but also raise critical ethical concerns. Most studies focus on narrow ethical dimensions, and also limited diversity of languages and models. To address the…

DiversityFairness

LLMGeo: Benchmarking Large Language Models on Image Geolocation In-the-wild

2024-05-30 · Zhiqiang Wang, Dejia Xu, Rana Muhammad Shahroz Khan, Yanbin Lin 외

Image geolocation is a critical task in various image-understanding applications. However, existing methods often fail when analyzing challenging, in-the-wild images. Inspired by the exceptional background knowledge of m…

Benchmarking

OpenJAI-v1.0: An Open Thai Large Language Model

2025-10-08 · Pontakorn Trakuekul, Attapol T. Rutherford, Jullajak Karnjanaekarin, Narongkorn Panitsrisit 외 arxiv

We introduce OpenJAI-v1.0, an open-source large language model for Thai and English, developed from the Qwen3-14B model. Our work focuses on boosting performance on practical tasks through carefully curated data across t…

Long-Context UnderstandingInstruction Following

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

2023-08-02 · Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel 외

We introduce OpenFlamingo, a family of autoregressive vision-language models ranging from 3B to 9B parameters. OpenFlamingo is an ongoing effort to produce an open-source replication of DeepMind's Flamingo models. On sev…

Visual Question AnsweringVisual Question Answering (VQA)