Large language model as user daily behavior data generator: balancing population diversity and individual personality
Predicting human daily behavior is challenging due to the complexity of routine patterns and short-term fluctuations. While data-driven models have improved behavior prediction by leveraging empirical data from various platforms and devices, the reliance on sensitive, large-scale user data raises privacy concerns and limits data availability. Synthetic data generation has emerged as a promising solution, though existing methods are often limited to specific applications. In this work, we introduce BehaviorGen, a framework that uses large language models (LLMs) to generate high-quality synthetic behavior data. By simulating user behavior based on profiles and real events, BehaviorGen supports data augmentation and replacement in behavior prediction models. We evaluate its performance in scenarios such as pertaining augmentation, fine-tuning replacement, and fine-tuning augmentation, achieving significant improvements in human mobility and smartphone usage predictions, with gains of up to 18.9%. Our results demonstrate the potential of BehaviorGen to enhance user behavior modeling through flexible and privacy-preserving synthetic data generation.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationDiversityLanguage ModelingLanguage ModellingLarge Language ModelPrivacy PreservingSynthetic Data GenerationSimilar Papers 제목 키워드 기반
Learning Language and Multimodal Privacy-Preserving Markers of Mood from Mobile Data
Mental health conditions remain underdiagnosed even in countries with common access to advanced medical care. The ability to accurately and efficiently predict mood from easily collectible data has several important impl…
Privacy PreservingMultimodal Privacy-preserving Mood Prediction from Mobile Data: A Preliminary Study
Mental health conditions remain under-diagnosed even in countries with common access to advanced medical care. The ability to accurately and efficiently predict mood from easily collectible data has several important imp…
Privacy PreservingLUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs
Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterog…
SeqLLM: Augmenting LLMs with Behavioral-Sequence Modeling for High-Stakes Decisions at WeChat Pay
Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false negatives leave harmful activity undetected. The hardest cases requi…
Breaking the Length Barrier: LLM-Enhanced CTR Prediction in Long Textual User Behaviors
With the rise of large language models (LLMs), recent works have leveraged LLMs to improve the performance of click-through rate (CTR) prediction. However, we argue that a critical obstacle remains in deploying LLMs for …
Click-Through Rate Prediction