paper-with-me

홈 › Papers

Can Smaller Large Language Models Evaluate Research Quality?

2025-08-10 · Mike Thelwall arxiv

Although both Google Gemini (1.5 Flash) and ChatGPT (4o and 4o-mini) give research quality evaluation scores that correlate positively with expert scores in nearly all fields, and more strongly that citations in most, it is not known whether this is true for smaller Large Language Models (LLMs). In response, this article assesses Google's Gemma-3-27b-it, a downloadable LLM (60Gb). The results for 104,187 articles show that Gemma-3-27b-it scores correlate positively with an expert research quality score proxy for all 34 Units of Assessment (broad fields) from the UK Research Excellence Framework 2021. The Gemma-3-27b-it correlations have 83.8% of the strength of ChatGPT 4o and 94.7% of the strength of ChatGPT 4o-mini correlations. Differently from the two larger LLMs, the Gemma-3-27b-it correlations do not increase substantially when the scores are averaged across five repetitions, its scores tend to be lower, and its reports are relatively uniform in style. Overall, the results show that research quality score estimation can be conducted by offline LLMs, so this capability is not an emergent property of the largest LLMs. Moreover, score improvement through repetition is not a universal feature of LLMs. In conclusion, although the largest LLMs still have the highest research evaluation score estimation capability, smaller ones can also be used for this task, and this can be helpful for cost saving or when secure offline processing is needed.

📄 PDF Abstract BibTeX arXiv:2508.07196

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Can Small and Reasoning Large Language Models Score Journal Articles for Research Quality and Do Averaging and Few-shot Help?

2025-10-25 · Mike Thelwall, Ehsan Mohammadi arxiv

Previous research has shown that journal article quality ratings from the cloud based Large Language Model (LLM) families ChatGPT and Gemini and the medium sized open weights LLM Gemma3 27b correlate moderately with expe…

CollectiveSFT: Scaling Large Language Models for Chinese Medical Benchmark with Collective Instructions in Healthcare

2024-07-29 · Jingwei Zhu, Minghuan Tan, Min Yang, Ruixue Li 외

The rapid progress in Large Language Models (LLMs) has prompted the creation of numerous benchmarks to evaluate their capabilities.This study focuses on the Comprehensive Medical Benchmark in Chinese (CMB), showcasing ho…

Diversity

Automatic generation of a large dictionary with concreteness/abstractness ratings based on a small human dictionary

2022-06-13 · Vladimir Ivanov, Valery Solovyev

Concrete/abstract words are used in a growing number of psychological and neurophysiological research. For a few languages, large dictionaries have been created manually. This is a very time-consuming and costly process.…

The Ideation Bottleneck: Decomposing the Quality Gap Between AI-Generated and Human Economics Research

2026-04-03 · Ning Li arxiv

Autonomous AI systems can now generate complete economics research papers, but they substantially underperform human-authored publications in head-to-head comparisons. This paper decomposes the quality gap into two indep…

Enhancing Scientific Discourse: Machine Translation for the Scientific Domain

2026-05-20 · Dimitris Roussis, Sokratis Sofianopoulos, Stelios Piperidis arxiv

The increasing volume of scientific research necessitates effective communication across language barriers. Machine translation (MT) offers a promising solution for accessing international publications. However, the scie…

Machine Translation