paper-with-me

홈 › Papers

Domain-based Latent Personal Analysis and its use for impersonation detection in social media

2020-04-05 · Osnat Mokryn, Hagit Ben-Shoshan

Zipf's law defines an inverse proportion between a word's ranking in a given corpus and its frequency in it, roughly dividing the vocabulary into frequent words and infrequent ones. Here, we stipulate that within a domain an author's signature can be derived from, in loose terms, the author's missing popular words and frequently used infrequent-words. We devise a method, termed Latent Personal Analysis (LPA), for finding domain-based attributes for entities in a domain: their distance from the domain and their signature, which determines how they most differ from a domain. We identify the most suitable distance metric for the method among several and construct the distances and personal signatures for authors, the domain's entities. The signature consists of both over-used terms (compared to the average), and missing popular terms. We validate the correctness and power of the signatures in identifying users and set existence conditions. We then show uses for the method in explainable authorship attribution: we define algorithms that utilize LPA to identify two types of impersonation in social media: (1) authors with sockpuppets (multiple) accounts; (2) front users accounts, operated by several authors. We validate the algorithms and employ them over a large scale dataset obtained from a social media site with over 4000 users. We corroborate these results using temporal rate analysis. LPA can further be used to devise personal attributes in a wide range of scientific domains in which the constituents have a long-tail distribution of elements.

📄 PDF Abstract BibTeX arXiv:2004.02346

Code (1)

ScanLab-ossi/LPA 공식 구현

Tasks

Authorship Attribution

Similar Papers 제목 키워드 기반

When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection

2025-10-14 · Lang Gao, Xuhui Li, Chenxi Wang, Mingzhe Li 외 arxiv

Large language models (LLMs) have grown more powerful in language generation, producing fluent text and even imitating personal style. Yet, this ability also heightens the risk of identity impersonation. To the best of o…

Text Detection

IMPersona: Evaluating Individual Level LM Impersonation

2025-04-06 · Quan Shi, Carlos E. Jimenez, Stephen Dong, Brian Seo 외

As language models achieve increasingly human-like capabilities in conversational text generation, a critical question emerges: to what extent can these systems simulate the characteristics of specific individuals? To ev…

Text Generation

GaitProtector: Impersonation-Driven Gait De-Identification via Training-Free Diffusion Latent Optimization

2026-05-12 · Huiran Duan, Qian Zhou, Zhongliang Guo, Junhao Dong 외 arxiv

Conventional gait de-identification methods often encounter an inherent trade-off: they either provide insufficient identity suppression or introduce spatiotemporal distortions that impede structure-sensitive downstream …

Gait Recognition

DIRF: A Framework for Digital Identity Protection and Clone Governance in Agentic AI Systems

2025-08-04 · Hammad Atta, Muhammad Zeeshan Baig, Yasir Mehmood, Nadeem Shahzad 외 arxiv

The rapid advancement and widespread adoption of generative artificial intelligence (AI) pose significant threats to the integrity of personal identity, including digital cloning, sophisticated impersonation, and the una…

Phantom: A Unified Face-Swap Deepfake Protection Framework with Latent and Spatial Constraints

2026-06-30 · Jungkon Kim, Cheolseung Jung, Jong-Min Choi, Juseong Lee arxiv

Face-swapping deepfakes pose an escalating threat to personal privacy by enabling unauthorized identity manipulation. While adversarial approaches have demonstrated success against black-box face recognition (FR) models,…

Face Recognition