paper-with-me

홈 › Papers

SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding

2026-01-19 · Xiaohan Huang, Meng Xiao, Chuan Qin, Qingqing Long, Jinmiao Chen, Yuanchun Zhou, Hengshu Zhu arxiv

Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks. However, their ability to reliably reason from gene-level knowledge to functional understanding, a core requirement for knowledge-enhanced cell atlas interpretation, remains largely underexplored. To address this gap, we introduce SciHorizon-GENE, a large-scale gene-centric benchmark constructed from authoritative biological databases. The benchmark integrates curated knowledge for over 190K human genes and comprises more than 540K questions covering diverse gene-to-function reasoning scenarios relevant to cell type annotation, functional interpretation, and mechanism-oriented analysis. Motivated by behavioral patterns observed in preliminary examinations, SciHorizon-GENE evaluates LLMs along four biologically critical perspectives: research attention sensitivity, hallucination tendency, answer completeness, and literature influence, explicitly targeting failure modes that limit the safe adoption of LLMs in biological interpretation pipelines. We systematically evaluate a wide range of state-of-the-art general-purpose and biomedical LLMs, revealing substantial heterogeneity in gene-level reasoning capabilities and persistent challenges in generating faithful, complete, and literature-grounded functional interpretations. Our benchmark establishes a systematic foundation for analyzing LLM behavior at the gene scale and offers insights for model selection and development, with direct relevance to knowledge-enhanced biological interpretation.

📄 PDF Abstract BibTeX arXiv:2601.12805

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models

2025-03-12 · Chuan Qin, Xin Chen, Chengrui Wang, Pengmin Wu 외

In recent years, the rapid advancement of Artificial Intelligence (AI) technologies, particularly Large Language Models (LLMs), has revolutionized the paradigm of scientific discovery, establishing AI-for-Science (AI4Sci…

BenchmarkingFairnessscientific discovery

SciHorizon-DataEVA: An Agentic System for AI-Readiness Evaluation of Heterogeneous Scientific Data

2026-04-29 · Dianyu Liu, Chuan Qin, Xi Chen, Xiaohan Li 외 arxiv

AI-for-Science (AI4Science) is increasingly transforming scientific discovery by embedding machine learning models into prediction, simulation, and hypothesis generation workflows across domains. However, the effectivene…

Four Shades of Life Sciences: A Dataset for Disinformation Detection in the Life Sciences

2025-07-04 · Eva Seidlmayer, Lukas Galke, Konrad U. Förstner arxiv

Disseminators of disinformation often seek to attract attention or evoke emotions - typically to gain influence or generate revenue - resulting in distinctive rhetorical patterns that can be exploited by machine learning…

Fact Checking

Fluorescent Neuronal Cells v2: Multi-Task, Multi-Format Annotations for Deep Learning in Microscopy

2023-07-26 · Luca Clissa, Antonio Macaluso, Roberto Morelli, Alessandra Occhinegro 외

Fluorescent Neuronal Cells v2 is a collection of fluorescence microscopy images and the corresponding ground-truth annotations, designed to foster innovative research in the domains of Life Sciences and Deep Learning. Th…

Benchmarkingobject-detectionObject DetectionSegmentation+3

Computing in the Life Sciences: From Early Algorithms to Modern AI

2024-06-17 · Samuel A. Donkor, Matthew E. Walsh, Alexander J. Titus

Computing in the life sciences has undergone a transformative evolution, from early computational models in the 1950s to the applications of artificial intelligence (AI) and machine learning (ML) seen today. This paper h…

Decision Making