paper-with-me

Papers

Identifying Non-Replicable Social Science Studies with Language Models

2025-03-10 · Denitsa Saynova, Kajsa Hansson, Bastiaan Bruinsma, Annika Fredén, Moa Johansson

In this study, we investigate whether LLMs can be used to indicate if a study in the behavioural social sciences is replicable. Using a dataset of 14 previously replicated studies (9 successful, 5 unsuccessful), we evaluate the ability of both open-source (Llama 3 8B, Qwen 2 7B, Mistral 7B) and proprietary (GPT-4o) instruction-tuned LLMs to discriminate between replicable and non-replicable findings. We use LLMs to generate synthetic samples of responses from behavioural studies and estimate whether the measured effects support the original findings. When compared with human replication results for these studies, we achieve F1 values of up to $77\%$ with Mistral 7B, $67\%$ with GPT-4o and Llama 3 8B, and $55\%$ with Qwen 2 7B, suggesting their potential for this task. We also analyse how effect size calculations are affected by sampling temperature and find that low variance (due to temperature) leads to biased effect estimates.

📄 PDF Abstract BibTeX arXiv:2503.10671

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

Replicable Reinforcement Learning

2023-05-24 · NeurIPS 2023 11

The replicability crisis in the social, behavioral, and data sciences has led to the formulation of algorithm frameworks for replicability -- i.e., a requirement that an algorithm produce identical outputs (with high pro…

reinforcement-learningReinforcement Learning

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences

2026-02-11 · Bang Nguyen, Dominik Soós, Qian Ma, Rochana R. Obadage 외 arxiv

The literature has witnessed an emerging interest in AI agents for automated assessment of scientific papers. Existing benchmarks focus primarily on the computational aspect of this task, testing agents' ability to repro…

A Financial Brain Scan of the LLM

2025-08-29 · Hui Chen, Antoine Didisheim, Mohammad, Pourmohammadi 외 arxiv

Emerging techniques in computer science make it possible to "brain scan" large language models (LLMs), identify the plain-English concepts that guide their reasoning, and steer them while holding other factors constant. …

From Labor to Collaboration: A Methodological Experiment Using AI Agents to Augment Research Perspectives in Taiwan's Humanities and Social Sciences

2026-02-19 · Yi-Chih Huang arxiv

Generative AI is reshaping knowledge work, yet existing research focuses predominantly on software engineering and the natural sciences, with limited methodological exploration for the humanities and social sciences. Pos…

Information RetrievalText Generation

Identifying and Characterizing Active Citizens who Refute Misinformation in Social Media

2022-04-21 · Yida Mu, Pu Niu, Nikolaos Aletras

The phenomenon of misinformation spreading in social media has developed a new form of active citizens who focus on tackling the problem by refuting posts that might contain misinformation. Automatically identifying and …

Misinformation