paper-with-me

홈 › Papers

Selecting Language Models for Social Science: Start Small, Start Open, and Validate

2026-01-16 · Dustin S. Stoltz, Marshall A. Taylor, Sanuj Kumar arxiv

Currently, there are thousands of large pretrained language models (LLMs) available to social scientists. How do we select among them? Using validity, reliability, reproducibility, and replicability as guides, we explore the significance of: (1) model openness, (2) model footprint, (3) training data, and (4) model architectures and fine-tuning. While ex-ante tests of validity (i.e., benchmarks) are often privileged in these discussions, we argue that social scientists cannot altogether avoid validating computational measures (ex-post). Replicability, in particular, is a more pressing guide for selecting language models. Being able to reliably replicate a particular finding that entails the use of a language model necessitates reliably reproducing a task. To this end, we propose starting with smaller, open models, and constructing delimited benchmarks to demonstrate the validity of the entire computational pipeline.

📄 PDF Abstract BibTeX arXiv:2601.10926

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluation of Word Embeddings for the Social Sciences

2023-02-13 · LaTeCHCLfL (COLING) 2022 10 · Ricardo Schiffers, Dagmar Kern, Daniel Hienert

Word embeddings are an essential instrument in many NLP tasks. Most available resources are trained on general language from Web corpora or Wikipedia dumps. However, word embeddings for domain-specific language are rare,…

DiversityWord Embeddings

Machine learning in the social and health sciences

2021-06-20 · Anja K. Leist, Matthias Klee, Jung Hyun Kim, David H. Rehkopf 외

The uptake of machine learning (ML) approaches in the social and health sciences has been rather slow, and research using ML for social and health research questions remains fragmented. This may be due to the separate de…

BIG-bench Machine LearningCausal Inference

Identifying and Improving Dataset References in Social Sciences Full Texts

2016-03-29 · Ghavimi Behnam, Mayr Philipp, Vahdati Sahar, Lange Christoph

Scientific full text papers are usually stored in separate places than their underlying research datasets. Authors typically make references to datasets by mentioning them for example by using their titles and the year o…

Inferring users' preferences through leveraging their social relationships

2017-11-28 · Deng Xiaofang, Wu Leilei, Ren Xiaolong, Jia Chunxiao 외

Recommender systems, inferring users' preferences from their historical activities and personal profiles, have been an enormous success in the last several years. Most of the existing works are based on the similarities …

Recommendation Systems

Towards Coding Social Science Datasets with Language Models

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Researchers often rely on humans to code (label, annotate, etc.) large sets of texts. This is a highly variable task and requires a great deal of time and resources. Efforts to automate this process have achieved human-l…