paper-with-me

홈 › Papers

BenCSSmark: Making the Social Sciences Count in LLM Research

2026-05-06 · Arnault Chatelain, Étienne Ollion, Qianwen Guan, Diandra Fabre, Lorraine Goeuriot, Emile Chapuis, Abdelkrim Beloued, Marie Candito, Nicolas Hervé, Didier Schwab arxiv

This position paper argues that the under-representation of social science tasks in contemporary LLM benchmarks limits advances in both LLM evaluation and social scientific inquiry. Benchmarks -- standardized tools for assessing computational systems -- are pivotal in the development of artificial intelligence (AI), including large language models (LLMs). Benchmarks do more than measure progress -- they actively structure it, shaping reputations, research agendas, and commercial outcomes. Despite this central role, the social sciences are largely absent from mainstream evaluation frameworks, even though scholars in these fields generate dozens of rigorously annotated, context-sensitive datasets each year. Integrating this work into benchmark design could significantly improve the generalization and robustness of AI models. In turn, models trained on social scientific tasks would likely yield better performance on classic and contemporary tasks in disciplines as diverse as history, sociology, political science or economics. This is all the more pressing as these disciplines are quickly turning to LLMs for assistance. To address this gap, we introduce BenCSSmark, a benchmark composed of datasets annotated by computational social scientists. By integrating social scientific perspectives into benchmarking, BenCSSmark seeks to promote more robust, transparent, and socially relevant AI systems and to foster efficient collaboration.

📄 PDF Abstract BibTeX arXiv:2605.04886

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Research Design Tracking and Assessment for the Social Sciences

2026-08-27 · Marco Rovera, Sergiu Burlacu, Dominique Cappelletti, Alessio Tomelleri 외 arxiv

Reliable assessment of causal research designs in the social sciences is critical for evidence-based policy-making, yet has so far relied entirely on manual expert analysis. We introduce Automated Research Design Trackin…

Machine learning in the social and health sciences

2021-06-20 · Anja K. Leist, Matthias Klee, Jung Hyun Kim, David H. Rehkopf 외

The uptake of machine learning (ML) approaches in the social and health sciences has been rather slow, and research using ML for social and health research questions remains fragmented. This may be due to the separate de…

BIG-bench Machine LearningCausal Inference

A bibliometric analysis and scoping study to identify English-language perspectives on slums

2024-12-18 · Katharina Henn, Michaela Lestakova, Kevin Logan, Jakob Hartig 외

Slums, informal settlements, and deprived areas are urban regions characterized by poverty. According to the United Nations, over one billion people reside in these areas, and this number is projected to increase. Additi…

Artificial Intelligence and work: a critical review of recent research from the social sciences

2022-02-25 · Jean-Philippe Deranty, Thomas Corbin

This review seeks to present a comprehensive picture of recent discussions in the social sciences of the anticipated impact of AI on the world of work. Issues covered include technological unemployment, algorithmic manag…

Management

Personalized Public Policy Analysis in Social Sciences using Causal-Graphical Normalizing Flows

2022-02-07 · Sourabh Balgi, Jose M. Pena, Adel Daoud

Structural Equation/Causal Models (SEMs/SCMs) are widely used in epidemiology and social sciences to identify and analyze the average causal effect (ACE) and conditional ACE (CACE). Traditional causal effect estimation m…

counterfactualCounterfactual InferenceEpidemiology