paper-with-me

Papers

SynDy: Synthetic Dynamic Dataset Generation Framework for Misinformation Tasks

2024-05-17 · Michael Shliselberg, Ashkan Kazemi, Scott A. Hale, Shiri Dori-Hacohen

Diaspora communities are disproportionately impacted by off-the-radar misinformation and often neglected by mainstream fact-checking efforts, creating a critical need to scale-up efforts of nascent fact-checking initiatives. In this paper we present SynDy, a framework for Synthetic Dynamic Dataset Generation to leverage the capabilities of the largest frontier Large Language Models (LLMs) to train local, specialized language models. To the best of our knowledge, SynDy is the first paper utilizing LLMs to create fine-grained synthetic labels for tasks of direct relevance to misinformation mitigation, namely Claim Matching, Topical Clustering, and Claim Relationship Classification. SynDy utilizes LLMs and social media queries to automatically generate distantly-supervised, topically-focused datasets with synthetic labels on these three tasks, providing essential tools to scale up human-led fact-checking at a fraction of the cost of human-annotated data. Training on SynDy's generated labels shows improvement over a standard baseline and is not significantly worse compared to training on human labels (which may be infeasible to acquire). SynDy is being integrated into Meedan's chatbot tiplines that are used by over 50 organizations, serve over 230K users annually, and automatically distribute human-written fact-checks via messaging apps such as WhatsApp. SynDy will also be integrated into our deployed Co-Insights toolkit, enabling low-resource organizations to launch tiplines for their communities. Finally, we envision SynDy enabling additional fact-checking tools such as matching new misinformation claims to high-quality explainers on common misinformation topics.

📄 PDF Abstract BibTeX arXiv:2405.10700

Code (0)

등록된 구현이 없습니다.

Tasks

ChatbotDataset GenerationFact CheckingMisinformation

Similar Papers 제목 키워드 기반

DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data

2026-04-02 · Wonjoon Jin, Jiyun Won, Janghyeok Han, Qi Dai 외 arxiv

Despite recent progress, video diffusion models still struggle to synthesize realistic videos involving highly dynamic motions or requiring fine-grained motion controllability. A central limitation lies in the scarcity o…

Equilibrium Dynamics and Mitigation of Gender Bias in Synthetically Generated Data

2025-11-12 · Ashish Kattamuri, Arpita Vats, Harshwardhan Fartale, Rahul Raja 외 arxiv

Recursive prompting with large language models enables scalable synthetic dataset generation but introduces the risk of bias amplification. We investigate gender bias dynamics across three generations of recursive text g…

Synthetic Data GenerationSemantic SimilarityText Generation

Energy-based Autoregressive Generation for Neural Population Dynamics

2025-11-18 · Ningling Ge, Sicheng Dai, Yu Zhu, Shan Yu arxiv

Understanding brain function represents a fundamental goal in neuroscience, with critical implications for therapeutic interventions and neural engineering applications. Computational modeling provides a quantitative fra…

Computational Efficiency

AC3S: Adaptive Conditioning for 3D-Aware Synthetic Data Generation

2026-06-30 · Eric Ji, Qiran Hu, Wufei Ma, Sarthak Jain 외 arxiv

Synthetic data generation has emerged as a powerful tool for improving data scalability in computer vision. Recent diffusion-based pipelines have demonstrated strong photorealism. However, how to enforce precise 3D struc…

Synthetic Data GenerationImage Generation

Synthesizing Diverse Network Flow Datasets with Scalable Dynamic Multigraph Generation

2025-05-12 · Arya Grayeli, Vipin Swarup, Steven E. Noel

Obtaining real-world network datasets is often challenging because of privacy, security, and computational constraints. In the absence of such datasets, graph generative models become essential tools for creating synthet…

DiversityGenerative Adversarial NetworkGraph Generation