paper-with-me

홈 › Papers

Unknown Aware AI-Generated Content Attribution

2026-01-01 · Ellie Thieu, Jifan Zhang, Haoyue Bai arxiv

The rapid advancement of photorealistic generative models has made it increasingly important to attribute the origin of synthetic content, moving beyond binary real or fake detection toward identifying the specific model that produced a given image. We study the problem of distinguishing outputs from a target generative model (e.g., OpenAI Dalle 3) from other sources, including real images and images generated by a wide range of alternative models. Using CLIP features and a simple linear classifier, shown to be effective in prior work, we establish a strong baseline for target generator attribution using only limited labeled data from the target model and a small number of known generators. However, this baseline struggles to generalize to harder, unseen, and newly released generators. To address this limitation, we propose a constrained optimization approach that leverages unlabeled wild data, consisting of images collected from the Internet that may include real images, outputs from unknown generators, or even samples from the target model itself. The proposed method encourages wild samples to be classified as non target while explicitly constraining performance on labeled data to remain high. Experimental results show that incorporating wild data substantially improves attribution performance on challenging unseen generators, demonstrating that unlabeled data from the wild can be effectively exploited to enhance AI generated content attribution in open world settings.

📄 PDF Abstract BibTeX arXiv:2601.00218

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sovereign Context Protocol: An Open Attribution Layer for Human-Generated Content in the Age of Large Language Models

2026-03-28 · Praneel Panchigar, Torlach Rush, Matthew Canabarro arxiv

Large Language Models (LLMs) consume vast quantities of human-generated content for both training and real-time inference, yet the creators of that content remain largely invisible in the value chain. Existing approaches…

Watermark-based Attribution of AI-Generated Content

2024-04-05 · Zhengyuan Jiang, Moyang Guo, Yuepeng Hu, Neil Zhenqiang Gong

Several companies have deployed watermark-based detection to identify AI-generated content. However, attribution--the ability to trace back to the user of a generative AI (GenAI) service who created a given piece of AI-g…

Progressive Open Space Expansion for Open-Set Model Attribution

2023-03-13 · CVPR 2023 1 · Tianyun Yang, Danding Wang, Fan Tang, Xinying Zhao 외

Despite the remarkable progress in generative technology, the Janus-faced issues of intellectual property protection and malicious content supervision have arisen. Efforts have been paid to manage synthetic images by att…

AttributeOpen Set Learning

SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling

2025-08-12 · Shixuan Sun, Siyuan Liang, Ruoyu Chen, Jianjie Huang 외 arxiv

Retrieval-Augmented Generation (RAG) and its Multimodal Retrieval-Augmented Generation (MRAG) significantly improve the knowledge coverage and contextual understanding of Large Language Models (LLMs) by introducing exter…

Image Retrieval

Who Owns the Output? Bridging Law and Technology in LLMs Attribution

2025-03-29 · Emanuele Mezzi, Asimina Mertzani, Michael P. Manis, Siyanna Lilova 외

Since the introduction of ChatGPT in 2022, Large language models (LLMs) and Large Multimodal Models (LMM) have transformed content creation, enabling the generation of human-quality content, spanning every medium, text, …