paper-with-me

Papers

Scaling Synthetic Data Creation with 1,000,000,000 Personas

2024-06-28 · Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, Dong Yu

We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce Persona Hub -- a collection of 1 billion diverse personas automatically curated from web data. These 1 billion personas (~13% of the world's total population), acting as distributed carriers of world knowledge, can tap into almost every perspective encapsulated within the LLM, thereby facilitating the creation of diverse synthetic data at scale for various scenarios. By showcasing Persona Hub's use cases in synthesizing high-quality mathematical and logical reasoning problems, instructions (i.e., user prompts), knowledge-rich texts, game NPCs and tools (functions) at scale, we demonstrate persona-driven data synthesis is versatile, scalable, flexible, and easy to use, potentially driving a paradigm shift in synthetic data creation and applications in practice, which may have a profound impact on LLM research and development.

📄 PDF Abstract BibTeX arXiv:2406.20094

Code (4)

tencent-ailab/persona-hub 공식 구현
camel-ai/camel pytorch
goodmike31/pl-asr-bigos-tools
lightaime/camel pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelLogical ReasoningWorld Knowledge

Similar Papers 제목 키워드 기반

The Need for a Socially-Grounded Persona Framework for User Simulation

2026-01-12 · Pranav Narayanan Venkit, Yu Li, Yada Pruksachatkun, Chien-Sheng Wu arxiv

Synthetic personas are widely used to condition large language models (LLMs) for social simulation, yet most personas are still constructed from coarse sociodemographic attributes or summaries. We revisit persona creatio…

DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas

2025-11-10 · Zhen Wang, Yufan Zhou, Zhongyan Luo, Lyumanshan Ye 외 arxiv

Simulating human profiles by instilling personas into large language models (LLMs) is rapidly transforming research in agentic behavioral simulation, LLM personalization, and human-AI alignment. However, most existing sy…

Question Answering

Modeling Human Subjectivity in LLMs Using Explicit and Implicit Human Factors in Personas

2024-06-20 · Salvatore Giorgi, Tingting Liu, Ankit Aich, Kelsey Isman 외

Large language models (LLMs) are increasingly being used in human-centered social scientific tasks, such as data annotation, synthetic data creation, and engaging in dialog. However, these tasks are highly subjective and…

Language ModellingLarge Language Model

SDialog: A Python Toolkit for Synthetic Dialogue Generation and Analysis

2025-06-12 · Sergio Burdisso, Esaú Villatoro-Tello, Petr Motlicek

The advancement of conversational AI systems relies on the availability of high-quality, flexible, and reproducible synthetic dialogues for training, evaluation, and benchmarking. SDialog is a modular, extensible Python …

BenchmarkingDialogue GenerationManagementSynthetic Data Generation

1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis

2024-06-17 · Sewade Ogun, Abraham T. Owodunni, Tobi Olatunji, Eniola Alese 외

Recent advances in speech synthesis have enabled many useful applications like audio directions in Google Maps, screen readers, and automated content generation on platforms like TikTok. However, these systems are mostly…

DiversitySpeech Synthesis