paper-with-me

홈 › Papers

Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions

2025-02-24 · Joseph Suh, Erfan Jahanparast, Suhong Moon, Minwoo Kang, Serina Chang

Large language models (LLMs) present novel opportunities in public opinion research by predicting survey responses in advance during the early stages of survey design. Prior methods steer LLMs via descriptions of subpopulations as LLMs' input prompt, yet such prompt engineering approaches have struggled to faithfully predict the distribution of survey responses from human subjects. In this work, we propose directly fine-tuning LLMs to predict response distributions by leveraging unique structural characteristics of survey data. To enable fine-tuning, we curate SubPOP, a significantly scaled dataset of 3,362 questions and 70K subpopulation-response pairs from well-established public opinion surveys. We show that fine-tuning on SubPOP greatly improves the match between LLM predictions and human responses across various subpopulations, reducing the LLM-human gap by up to 46% compared to baselines, and achieves strong generalization to unseen surveys and subpopulations. Our findings highlight the potential of survey-based fine-tuning to improve opinion prediction for diverse, real-world subpopulations and therefore enable more efficient survey designs. Our code is available at https://github.com/JosephJeesungSuh/subpop.

📄 PDF Abstract BibTeX arXiv:2502.16761

Code (1)

josephjeesungsuh/subpop 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingPrompt EngineeringSurvey

Similar Papers 제목 키워드 기반

Scaled Prompt-Tuning for Few-Shot Natural Language Generation

2023-09-13 · Ting Hu, Christoph Meinel, Haojin Yang

The increasingly Large Language Models (LLMs) demonstrate stronger language understanding and generation capabilities, while the memory demand and computation cost of fine-tuning LLMs on downstream tasks are non-negligib…

parameter-efficient fine-tuningText Generation

Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism

2025-10-15 · Xiaoshu Chen, Sihang Zhou, Ke Liang, Duanyang Yuan 외 arxiv

Chain of thought (CoT) fine-tuning aims to endow large language models (LLMs) with reasoning capabilities by training them on curated reasoning traces. It leverages both supervised and reinforced fine-tuning to cultivate…

Mathematical ReasoningCode Generation

The Interplay of Variant, Size, and Task Type in Arabic Pre-trained Language Models

2021-03-11 · EACL (WANLP) 2021 4 · Go Inoue, Bashar Alhafni, Nurpeiis Baimukan, Houda Bouamor 외

In this paper, we explore the effects of language variants, data sizes, and fine-tuning task types in Arabic pre-trained language models. To do so, we build three pre-trained language models across three variants of Arab…

Language ModelingLanguage Modelling

On the Crucial Role of Initialization for Matrix Factorization

2024-10-24 · Bingcong Li, Liang Zhang, Aryan Mokhtari, Niao He

This work revisits the classical low-rank matrix factorization problem and unveils the critical role of initialization in shaping convergence rates for such nonconvex and nonsmooth optimization. We introduce Nystrom init…

Scaling Up Biomedical Vision-Language Models: Fine-Tuning, Instruction Tuning, and Multi-Modal Learning

2025-05-23 · Cheng Peng, Kai Zhang, Mengxian Lyu, Hongfang Liu 외

To advance biomedical vison-language model capabilities through scaling up, fine-tuning, and instruction tuning, develop vision-language models with improved performance in handling long text, explore strategies to effic…

DecoderImage Captioningimage-classificationImage Classification+6