paper-with-me

홈 › Papers

Protecting Vulnerable Voices: Synthetic Dataset Generation for Self-Disclosure Detection

2025-07-24 · Shalini Jangra, Suparna De, Nishanth Sastry, Saeed Fadaei arxiv

Social platforms such as Reddit have a network of communities of shared interests, with a prevalence of posts and comments from which one can infer users' Personal Information Identifiers (PIIs). While such self-disclosures can lead to rewarding social interactions, they pose privacy risks and the threat of online harms. Research into the identification and retrieval of such risky self-disclosures of PIIs is hampered by the lack of open-source labeled datasets. To foster reproducible research into PII-revealing text detection, we develop a novel methodology to create synthetic equivalents of PII-revealing data that can be safely shared. Our contributions include creating a taxonomy of 19 PII-revealing categories for vulnerable populations and the creation and release of a synthetic PII-labeled multi-text span dataset generated from 3 text generation Large Language Models (LLMs), Llama2-7B, Llama3-8B, and zephyr-7b-beta, with sequential instruction prompting to resemble the original Reddit posts. The utility of our methodology to generate this synthetic dataset is evaluated with three metrics: First, we require reproducibility equivalence, i.e., results from training a model on the synthetic data should be comparable to those obtained by training the same models on the original posts. Second, we require that the synthetic data be unlinkable to the original users, through common mechanisms such as Google Search. Third, we wish to ensure that the synthetic data be indistinguishable from the original, i.e., trained humans should not be able to tell them apart. We release our dataset and code at https://netsys.surrey.ac.uk/datasets/synthetic-self-disclosure/ to foster reproducible research into PII privacy risks in online social media.

📄 PDF Abstract BibTeX arXiv:2507.22930

Code (0)

등록된 구현이 없습니다.

Tasks

Text GenerationText Detection

Similar Papers 제목 키워드 기반

Protecting Your Voice: Temporal-aware Robust Watermarking

2025-04-21 · Yue Li, Weizhi Liu, Dongdong Lin, Hui Tian 외

The rapid advancement of generative models has led to the synthesis of real-fake ambiguous voices. To erase the ambiguity, embedding watermarks into the frequency-domain features of synthesized voices has become a common…

Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model

2024-07-24 · Jan Lehečka, Zdeněk Hanzlíček, Jindřich Matoušek, Daniel Tihelka

In this paper, we experimented with the SpeechT5 model pre-trained on large-scale datasets. We pre-trained the foundation model from scratch and fine-tuned it on a large-scale robust multi-speaker text-to-speech (TTS) ta…

text-to-speechText to Speech

Improved Child Text-to-Speech Synthesis through Fastpitch-based Transfer Learning

2023-11-07 · Rishabh Jain, Peter Corcoran

Speech synthesis technology has witnessed significant advancements in recent years, enabling the creation of natural and expressive synthetic speech. One area of particular interest is the generation of synthetic child s…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+5

Synthetic Data Generation Techniques for Developing AI-based Speech Assessments for Parkinson's Disease (A Comparative Study)

2023-12-04 · Mahboobeh Parsapoor

Changes in speech and language are among the first signs of Parkinson's disease (PD). Thus, clinicians have tried to identify individuals with PD from their voices for years. Doctors can leverage AI-based speech assessme…

Synthetic Data Generation

Adversarial speech for voice privacy protection from Personalized Speech generation

2024-01-22 · Shihao Chen, Liping Chen, Jie Zhang, KongAik Lee 외

The rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human list…

Speaker Verificationtext-to-speechText to SpeechVoice Conversion