paper-with-me

Papers

Large language model as user daily behavior data generator: balancing population diversity and individual personality

2025-05-23 · Haoxin Li, Jingtao Ding, Jiahui Gong, Yong Li

Predicting human daily behavior is challenging due to the complexity of routine patterns and short-term fluctuations. While data-driven models have improved behavior prediction by leveraging empirical data from various platforms and devices, the reliance on sensitive, large-scale user data raises privacy concerns and limits data availability. Synthetic data generation has emerged as a promising solution, though existing methods are often limited to specific applications. In this work, we introduce BehaviorGen, a framework that uses large language models (LLMs) to generate high-quality synthetic behavior data. By simulating user behavior based on profiles and real events, BehaviorGen supports data augmentation and replacement in behavior prediction models. We evaluate its performance in scenarios such as pertaining augmentation, fine-tuning replacement, and fine-tuning augmentation, achieving significant improvements in human mobility and smartphone usage predictions, with gains of up to 18.9%. Our results demonstrate the potential of BehaviorGen to enhance user behavior modeling through flexible and privacy-preserving synthetic data generation.

📄 PDF Abstract BibTeX arXiv:2505.17615

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationDiversityLanguage ModelingLanguage ModellingLarge Language ModelPrivacy PreservingSynthetic Data Generation

Similar Papers 제목 키워드 기반

Learning Language and Multimodal Privacy-Preserving Markers of Mood from Mobile Data

2021-06-24 · ACL 2021 5 · Paul Pu Liang, Terrance Liu, Anna Cai, Michal Muszynski 외

Mental health conditions remain underdiagnosed even in countries with common access to advanced medical care. The ability to accurately and efficiently predict mood from easily collectible data has several important impl…

Privacy Preserving

Multimodal Privacy-preserving Mood Prediction from Mobile Data: A Preliminary Study

2020-12-04 · Terrance Liu, Paul Pu Liang, Michal Muszynski, Ryo Ishii 외

Mental health conditions remain under-diagnosed even in countries with common access to advanced medical care. The ability to accurately and efficiently predict mood from easily collectible data has several important imp…

Privacy Preserving

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

2026-08-05 · Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao 외 arxiv

Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterog…

SeqLLM: Augmenting LLMs with Behavioral-Sequence Modeling for High-Stakes Decisions at WeChat Pay

2026-08-04 · Guilin Li, Jiaxing Zhang, Matthias Hwai Yong Tan, Bo Wang 외 arxiv

Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false negatives leave harmful activity undetected. The hardest cases requi…

Breaking the Length Barrier: LLM-Enhanced CTR Prediction in Long Textual User Behaviors

2024-03-28 · Binzong Geng, ZhaoXin Huan, Xiaolu Zhang, Yong He 외

With the rise of large language models (LLMs), recent works have leveraged LLMs to improve the performance of click-through rate (CTR) prediction. However, we argue that a critical obstacle remains in deploying LLMs for …

Click-Through Rate Prediction