paper-with-me

홈 › Papers

Remembering Unequally: Global and Disciplinary Bias in LLM Reconstruction of Scholarly Coauthor Lists

2025-11-01 · Ghazal Kalhor, Afra Mashhadi arxiv

Ongoing breakthroughs in large language models (LLMs) are reshaping scholarly search and discovery interfaces. While these systems offer new possibilities for navigating scientific knowledge, they also raise concerns about fairness and representational bias rooted in the models' memorized training data. As LLMs are increasingly used to answer queries about researchers and research communities, their ability to accurately reconstruct scholarly coauthor lists becomes an important but underexamined issue. In this study, we investigate how memorization in LLMs affects the reconstruction of coauthor lists and whether this process reflects existing inequalities across academic disciplines and world regions. We evaluate three prominent models, DeepSeek R1, Llama 4 Scout, and Mixtral 8x7B, by comparing their generated coauthor lists against bibliographic reference data. Our analysis reveals a systematic advantage for highly cited researchers, indicating that LLM memorization disproportionately favors already visible scholars. However, this pattern is not uniform: certain disciplines, such as Clinical Medicine, and some regions, including parts of Africa, exhibit more balanced reconstruction outcomes. These findings highlight both the risks and limitations of relying on LLM-generated relational knowledge in scholarly discovery contexts and emphasize the need for careful auditing of memorization-driven biases in LLM-based systems.

📄 PDF Abstract BibTeX arXiv:2511.00476

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Regression Conformal Prediction under Bias

2024-10-07 · Matt Y. Cheung, Tucker J. Netherton, Laurence E. Court, Ashok Veeraraghavan 외

Uncertainty quantification is crucial to account for the imperfect predictions of machine learning algorithms for high-impact applications. Conformal prediction (CP) is a powerful framework for uncertainty quantification…

Computed Tomography (CT)Conformal PredictionCT ReconstructionPrediction+4

Gender Bias in Contextualized Word Embeddings

2019-04-05 · NAACL 2019 6 · Jieyu Zhao, Tianlu Wang, Mark Yatskar, Ryan Cotterell 외

In this paper, we quantify, analyze and mitigate gender bias exhibited in ELMo's contextualized word vectors. First, we conduct several intrinsic analyses and find that (1) training data for ELMo contains significantly m…

Word Embeddings

Bias and Discrimination in AI: a cross-disciplinary perspective

2020-08-11 · Xavier Ferrer, Tom van Nuenen, Jose M. Such, Mark Coté 외

With the widespread and pervasive use of Artificial Intelligence (AI) for automated decision-making systems, AI bias is becoming more apparent and problematic. One of its negative consequences is discrimination: the unfa…

Decision Making

Geographic and Geopolitical Biases of Language Models

2022-12-20 · Fahim Faisal, Antonios Anastasopoulos

Pretrained language models (PLMs) often fail to fairly represent target users from certain world regions because of the under-representation of those regions in training datasets. With recent PLMs trained on enormous dat…

Complexity Reduction in the Negotiation of New Lexical Conventions

2018-05-15 · William Schueller, Vittorio Loreto, Pierre-Yves Oudeyer

In the process of collectively inventing new words for new concepts in a population, conflicts can quickly become numerous, in the form of synonymy and homonymy. Remembering all of them could cost too much memory, and re…