paper-with-me

Papers

SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models

2024-06-13 · Kehua Feng, Keyan Ding, Weijie Wang, Xiang Zhuang, Zeyuan Wang, Ming Qin, Yu Zhao, Jianhua Yao, Qiang Zhang, Huajun Chen

Large language models (LLMs) have gained increasing prominence in scientific research, but there is a lack of comprehensive benchmarks to fully evaluate their proficiency in understanding and mastering scientific knowledge. To address this need, we introduce the SciKnowEval benchmark, a novel framework that systematically evaluates LLMs across five progressive levels of scientific knowledge: studying extensively, inquiring earnestly, thinking profoundly, discerning clearly, and practicing assiduously. These levels aim to assess the breadth and depth of scientific knowledge in LLMs, including memory, comprehension, reasoning, discernment, and application. Specifically, we first construct a large-scale evaluation dataset encompassing 70K multi-level scientific problems and solutions in the domains of biology, chemistry, physics, and materials science. By leveraging this dataset, we benchmark 26 advanced open-source and proprietary LLMs using zero-shot and few-shot prompting strategies. The results reveal that despite the state-of-the-art performance of proprietary LLMs, there is still significant room for improvement, particularly in addressing scientific reasoning and applications. We anticipate that SciKnowEval will establish a standard for benchmarking LLMs in science research and promote the development of stronger scientific LLMs. The dataset and code are publicly available at https://scimind.ai/sciknoweval .

📄 PDF Abstract BibTeX arXiv:2406.09098

Code (1)

hicai-zju/sciknoweval 공식 구현

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models

2026-07-01 · Ye Liu, Srijan Bansal, Bo Pang, Yang Li 외 arxiv

Reinforcement learning with verifiable rewards (RLVR), along with recent selfdistillation variants such as SDPO, evaluates each rollout against a verifier and updates the policy from that episode-level signal. However, t…

Reinforcement Learning

Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning

2025-06-12 · Yuhao Zhou, Yiheng Wang, Xuming He, Ruoyao Xiao 외

Scientific discoveries increasingly rely on complex multimodal reasoning based on information-intensive scientific data and domain-specific expertise. Empowered by expert-level scientific benchmarks, scientific Multimoda…

AttributeMultimodal ReasoningVisual Question Answering (VQA)

I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization

2026-08-13 · Yubo Zhang, Xinhong Ma, Zezhong Tan, Ziqiang Dong arxiv

Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect. Privileged self-distillation can fill this…

Personalized Summarization of Scientific Scholarly Texts

2023-06-16 · Alka Khurana, Vasudha Bhatnagar, Vikas Kumar

In this paper, we present a proposal for an unsupervised algorithm, P-Summ, that generates an extractive summary of scientific scholarly text to meet the personal knowledge needs of the user. The method delves into the l…

ArticlesSentence

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

2026-08-10 · Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang 외 hf

We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across fou…

Video Generation