paper-with-me

홈 › Papers

Alignment Imprint: Zero-Shot AI-Generated Text Detection via Provable Preference Discrepancy

2026-04-18 · Junxi Wu, Kailin Huang, Dongjian Hu, Bin Chen, Hao Wu, Shu-Tao Xia, Changliang Zou arxiv

Detecting AI-generated text is an important but challenging problem. Existing likelihood-based detection methods are often sensitive to content complexity and may exhibit unstable performance. In this paper, our key insight is that modern Large Language Models (LLMs) undergo alignment (including fine-tuning and preference tuning), leaving a measurable distributional imprint. We theoretically derive this imprint by abstracting the alignment process as a sequence of constrained optimization steps, showing that the log-likelihood ratio can naturally decompose into implicit instructional biases and preference rewards. We refer to this quantity as the Alignment Imprint. Furthermore, to mitigate the instability in high-entropy regions, we introduce Log-likelihood Alignment Preference Discrepancy (LAPD), a standardized information-weighted statistic based on alignment imprint. We provide statistical guarantee that alignment-based statistics dominate Fast-DetectGPT in performance. We also theoretically show that LAPD strictly improves the unweighted alignment scores when the aligned and base models are close in distribution. Extensive experiments show that LAPD achieves an improvement 45.82% relative to the strongest existing baselines, yielding large and consistent gains across all settings.

📄 PDF Abstract BibTeX arXiv:2604.16923

Code (0)

등록된 구현이 없습니다.

Tasks

Text Detection

Similar Papers 제목 키워드 기반

An Empirical Investigation of Word Alignment Supervision for Zero-Shot Multilingual Neural Machine Translation

2021-11-01 · EMNLP 2021 11 · Alessandro Raganato, Raúl Vázquez, Mathias Creutz, Jörg Tiedemann

Zero-shot translations is a fascinating feature of Multilingual Neural Machine Translation (MNMT) systems. These MNMT models are usually trained on English-centric data, i.e. English either as the source or target langua…

Machine TranslationTranslationWord Alignment

Mitigating Hallucinations in Zero-Shot Scientific Summarisation: A Pilot Study

2025-11-30 · Imane Jaaouine, Ross D. King arxiv

Large language models (LLMs) produce context inconsistency hallucinations, which are LLM generated outputs that are misaligned with the user prompt. This research project investigates whether prompt engineering (PE) meth…

Prompt Engineering

Low-Shot Learning with Imprinted Weights

2017-12-19 · CVPR 2018 6 · Hang Qi, Matthew Brown, David G. Lowe

Human vision is able to immediately recognize novel visual categories after seeing just one or a few training examples. We describe how to add a similar capability to ConvNet classifiers by directly setting the final lay…

Revealing the Implicit Noise-based Imprint of Generative Models

2025-03-12 · Xinghan Li, Jingjing Chen, Yue Yu, Xue Song 외

With the rapid advancement of vision generation models, the potential security risks stemming from synthetic visual content have garnered increasing attention, posing significant challenges for AI-generated image detecti…

TriShield: Zero-Utility-Loss Defense Against Privacy Backdoors in Federated Language Model Fine-Tuning via Orthogonal Gradient Projection and Optimizer State Entanglement

2026-07-30 · Cheng Wei arxiv

Federated fine-tuning of large language models (LLMs) enables collaborative training without exposing raw data. However, a recent attack, NeuroImprint [1] (arXiv:2606.20553), demonstrates that a malicious parameter serve…