paper-with-me

홈 › Papers

InstaSynth: Opportunities and Challenges in Generating Synthetic Instagram Data with ChatGPT for Sponsored Content Detection

2024-03-22 · Thales Bertaglia, Lily Heisig, Rishabh Kaushal, Adriana Iamnitchi

Large Language Models (LLMs) raise concerns about lowering the cost of generating texts that could be used for unethical or illegal purposes, especially on social media. This paper investigates the promise of such models to help enforce legal requirements related to the disclosure of sponsored content online. We investigate the use of LLMs for generating synthetic Instagram captions with two objectives: The first objective (fidelity) is to produce realistic synthetic datasets. For this, we implement content-level and network-level metrics to assess whether synthetic captions are realistic. The second objective (utility) is to create synthetic data that is useful for sponsored content detection. For this, we evaluate the effectiveness of the generated synthetic data for training classifiers to identify undisclosed advertisements on Instagram. Our investigations show that the objectives of fidelity and utility may conflict and that prompt engineering is a useful but insufficient strategy. Additionally, we find that while individual synthetic posts may appear realistic, collectively they lack diversity, topic connectivity, and realistic user interaction patterns.

📄 PDF Abstract BibTeX arXiv:2403.15214

Code (1)

thalesbertaglia/instasynth 공식 구현

Tasks

DiversityPrompt Engineering

Similar Papers 제목 키워드 기반

Beyond Privacy: Navigating the Opportunities and Challenges of Synthetic Data

2023-04-07 · Boris van Breugel, Mihaela van der Schaar

Generating synthetic data through generative models is gaining interest in the ML community and beyond. In the past, synthetic data was often regarded as a means to private data release, but a surge of recent papers expl…

Data Augmentation

Detecting Calls to Action in Multimodal Content: Analysis of the 2021 German Federal Election Campaign on Instagram

2024-09-04 · Michael Achmann-Denkler, Jakob Fehle, Mario Haim, Christian Wolff

This study investigates the automated classification of Calls to Action (CTAs) within the 2021 German Instagram election campaign to advance the understanding of mobilization in social media contexts. We analyzed over 2,…

Robust classification

Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs

2025-07-24 · Tevin Atwal, Chan Nam Tieu, Yefeng Yuan, Zhan Shi 외 arxiv

The increasing use of synthetic data generated by Large Language Models (LLMs) presents both opportunities and challenges in data-driven applications. While synthetic data provides a cost-effective, scalable alternative …

Generative AI like ChatGPT in Blockchain Federated Learning: use cases, opportunities and future

2024-07-25 · Sai Puppala, Ismail Hossain, Md Jahangir Alam, Sajedul Talukder 외

Federated learning has become a significant approach for training machine learning models using decentralized data without necessitating the sharing of this data. Recently, the incorporation of generative artificial inte…

Federated Learning

Multilingual offensive lexicon annotated with contextual information

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Online hate speech and offensive comments detection is not a trivial research problem since pragmatic (contextual) factors influence what is considered offensive. Moreover, offensive terms are hardly found in classical l…

Abusive Language