paper-with-me

Papers

MetaBench: A Multi-task Benchmark for Assessing LLMs in Metabolomics

2025-10-16 · Yuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi, Rui Peng, Shuang Zeng, Xingyu Hu, Jinzhuo Wang, May D. Wang arxiv

Large Language Models (LLMs) have demonstrated remarkable capabilities on general text; however, their proficiency in specialized scientific domains that require deep, interconnected knowledge remains largely uncharacterized. Metabolomics presents unique challenges with its complex biochemical pathways, heterogeneous identifier systems, and fragmented databases. To systematically evaluate LLM capabilities in this domain, we introduce MetaBench, the first benchmark for metabolomics assessment. Curated from authoritative public resources, MetaBench evaluates five capabilities essential for metabolomics research: knowledge, understanding, grounding, reasoning, and research. Our evaluation of 25 open- and closed-source LLMs reveals distinct performance patterns across metabolomics tasks: while models perform well on text generation tasks, cross-database identifier grounding remains challenging even with retrieval augmentation. Model performance also decreases on long-tail metabolites with sparse annotations. With MetaBench, we provide essential infrastructure for developing and evaluating metabolomics AI systems, enabling systematic progress toward reliable computational tools for metabolomics research.

📄 PDF Abstract BibTeX arXiv:2510.14944

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

$\texttt{metabench}$ -- A Sparse Benchmark to Measure General Ability in Large Language Models

2024-07-04 · Alex Kipnis, Konstantinos Voudouris, Luca M. Schulze Buschoff, Eric Schulz

Large Language Models (LLMs) vary in their abilities on a range of tasks. Initiatives such as the $\texttt{Open LLM Leaderboard}$ aim to quantify these differences with several large benchmarks (sets of test items to whi…

ARCGSM8KHellaSwagMMLU+2

ArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language Models

2024-08-15 · Faris Hijazi, Somayah AlHarbi, Abdulaziz AlHussein, Harethah Abu Shairah 외

The rapid advancements in Large Language Models (LLMs) have led to significant improvements in various natural language processing tasks. However, the evaluation of LLMs' legal knowledge, particularly in non-English lang…

In-Context LearningMMLU

MetaGen: A DSL, Database, and Benchmark for VLM-Assisted Metamaterial Generation

2025-08-25 · Liane Makatura, Benjamin Jones, Siyuan Bian, Wojciech Matusik arxiv

Metamaterials are micro-architected structures whose geometry imparts highly tunable-often counter-intuitive-bulk properties. Yet their design is difficult because of geometric complexity and a non-trivial mapping from a…

AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

2025-01-03 · Sanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta, Yaoting Wang 외

With the rapid advancement of Multi-modal Large Language Models (MLLMs), several diagnostic benchmarks have recently been developed to assess these models' multi-modal reasoning proficiency. However, these benchmarks are…

Adversarial AttackDiagnostic

ArgBench: Benchmarking LLMs on Computational Argumentation Tasks

2026-04-19 · Yamen Ajjour, Carlotta Quensel, Nedim Lipka, Henning Wachsmuth arxiv

Argumentation skills are an essential toolkit for large language models (LLMs). These skills are crucial in various use cases, including self-reflection, debating collaboratively for diverse answers, and countering hate …