Synthetic Data Generation for Screen Time and App Usage
Smartphone usage data can provide valuable insights for understanding interaction with technology and human behavior. However, collecting large-scale, in-the-wild smartphone usage logs is challenging due to high costs, privacy concerns, under representative user samples and biases like non-response that can skew results. These challenges call for exploring alternative approaches to obtain smartphone usage datasets. In this context, large language models (LLMs) such as Open AI's ChatGPT present a novel approach for synthetic smartphone usage data generation, addressing limitations of real-world data collection. We describe a case study on how four prompt strategies influenced the quality of generated smartphone usage data. We contribute with insights on prompt design and measures of data quality, reporting a prompting strategy comparison combining two factors, prompt level of detail (describing a user persona, describing the expected results characteristics) and seed data inclusion (with versus without an initial real usage example). Our findings suggest that using LLMs to generate structured and behaviorally plausible smartphone use datasets is feasible for some use cases, especially when using detailed prompts. Challenges remain in capturing diverse nuances of human behavioral patterns in a single synthetic dataset, and evaluating tradeoffs between data fidelity and diversity, suggesting the need for use-case-specific evaluation metrics and future research with more diverse seed data and different LLM models.
Code (0)
등록된 구현이 없습니다.
Tasks
Synthetic Data GenerationSimilar Papers 제목 키워드 기반
BiometricBlender: Ultra-high dimensional, multi-class synthetic data generator to imitate biometric feature space
The lack of freely available (real-life or synthetic) high or ultra-high dimensional, multi-class datasets may hamper the rapidly growing research on feature screening, especially in the field of biometrics, where the us…
Synthetic Defect Generation for Display Front-of-Screen Quality Inspection: A Survey
Display front-of-screen (FOS) quality inspection is essential for the mass production of displays in the manufacturing process. However, the severe imbalanced data, especially the limited number of defect samples, has be…
Synthetic Data GenerationAnalysis of Benford’s Law for No-Reference Quality Assessment of Natural, Screen-Content, and Synthetic Images
With the tremendous growth and usage of digital images, no-reference image quality assessment is becoming increasingly important. This paper presents in-depth analysis of Benford’s law inspired first digit distribution f…
Image ForensicsImage Quality AssessmentNo-Reference Image Quality AssessmentSmall Target Detection for Search and Rescue Operations using Distributed Deep Learning and Synthetic Data Generation
It is important to find the target as soon as possible for search and rescue operations. Surveillance camera systems and unmanned aerial vehicles (UAVs) are used to support search and rescue. Automatic object detection i…
Data AugmentationImage Segmentationobject-detectionObject Detection+2Bayesian-Guided Generation of Synthetic Microbiomes with Minimized Pathogenicity
Synthetic microbiomes offer new possibilities for modulating microbiota, to address the barriers in multidtug resistance (MDR) research. We present a Bayesian optimization approach to enable efficient searching over the …
Bayesian OptimizationThompson Sampling