paper-with-me

홈 › Papers

LLM-Metrics: Measuring Research Impact Through Large Language Model Memory

2026-05-21 · Si Shen, Wenhua Zhao, Danhao Zhu arxiv

Citation counts remain the dominant metric for assessing research impact, yet they suffer from well-documented limitations: temporal lag, disciplinary bias, and Matthew effects. Here we propose LLM-Metrics, a research-impact assessment metric derived from the parametric memory of large language models (LLMs). The central hypothesis is that high-impact papers receive greater exposure in the academic community, that this exposure enters LLM training data in textual form, and that models consequently form stronger parametric memory of these papers. We designed four types of multiple-choice probes, covering title recognition, author recognition, method recognition, and venue recognition, and evaluated 549 computer science papers published in 2023-2024 across 17 LLMs spanning 0.5B to 72B parameters from six vendors. Of the 17 models, 15 produced positive predictions, 9 of which were significant at p less than 0.05, with an overall Spearman correlation of rho = 0.1495 and p = 0.0004 against citation counts. Three additional findings support the proposed mechanism. First, the predictive signal was stronger for 2024 papers, rho = 0.1880, whose citation counts were near zero at model-training time, reducing the plausibility of a simple reverse-causality explanation. Second, author-recognition probes showed the strongest discriminative power, consistent with an exposure-driven memory mechanism. Third, model scale and predictive power were non-monotonic: a 3B-parameter model, Llama-3.2-3B-Instruct, with rho = 0.1829, outperformed most larger models, supporting a selective-memory hypothesis in which the limited capacity of smaller models can serve as an effective information filter. LLM-Metrics offers a real-time, cross-disciplinary, citation-independent paradigm for research assessment.

📄 PDF Abstract BibTeX arXiv:2605.22176

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A global analysis of metrics used for measuring performance in natural language processing

2022-04-25 · nlppower (ACL) 2022 5 · Kathrin Blagec, Georg Dorffner, Milad Moradi, Simon Ott 외

Measuring the performance of natural language processing models is challenging. Traditionally used metrics, such as BLEU and ROUGE, originally devised for machine translation and summarization, have been shown to suffer …

BenchmarkingMachine Translation

Measuring prominence of scientific work in online news as a proxy for impact

2020-07-28 · James Ravenscroft, Amanda Clare, Maria Liakata

The impact made by a scientific paper on the work of other academics has many established metrics, including metrics based on citation counts and social media commenting. However, determination of the impact of a scienti…

ArticlesSemantic SimilaritySemantic Textual Similarity

The Impact of Research Funding on Knowledge Creation and Dissemination: A study of SNSF Research Grants

2020-11-23 · Rachel Heyard, Hanna Hottenrott

This study investigates the impact of competitive project-funding on researchers' publication outputs. Using detailed information on applicants at the Swiss National Science Foundation (SNSF) and their proposals' evaluat…

Articles

Measuring Ethics in AI with AI: A Methodology and Dataset Construction

2021-07-26 · Pedro H. C. Avelar, Rafael B. Audibert, Anderson R. Tavares, Luís C. Lamb

Recently, the use of sound measures and metrics in Artificial Intelligence has become the subject of interest of academia, government, and industry. Efforts towards measuring different phenomena have gained traction in t…

EthicsFairness

Optimising for Energy Efficiency and Performance in Machine Learning

2026-01-13 · Emile Dos Santos Ferreira, Andrei Paleyes, Neil D. Lawrence arxiv

The ubiquity of machine learning (ML) and the demand for ever-larger models bring an increase in energy consumption and environmental impact. However, little is known about the energy scaling laws in ML, and existing res…

Text Generation