paper-with-me

홈 › Papers

On the Salience of Low-Probability Tokens for AI-Generated Text Detection: A Multiscale Uncertainty Perspective

2026-06-01 · Yikai Guo, Bin Wang, Xilai Fan, Wenjun Ke, Haoran Luo arxiv

AI-generated text increasingly blends with human writing, raising practical risks such as misinformation, academic misuse, and corpora contamination. While statistical detectors are appealing for efficiency and generalization, they suffer from two key limitations. (i) Boilerplate dominance, boilerplate tokens shared across human and LLM writing can overwhelm discriminative signals. (ii) Brittle point estimates, relying on a single probability score yields unstable decisions under adversarial manipulations. To address these issues, we propose Uncertainty, a multiscale uncertainty estimator that focuses on informative low-probability tokens, which more clearly expose distributional discrepancies. Locally, it alleviates boilerplate dominance by averaging the log-probabilities of low-probability tokens; globally, it reduces brittleness by capturing the distributional shape of this low-probability region via Rényi entropy. We further extend the detector to Uncertainty++ via conditional independent sampling, yielding a more stable uncertainty estimation. Experiments across seven datasets and sixteen LLMs demonstrate high effectiveness, generalization, and robustness. Our code is available at https://github.com/guoyikai2000/Uncertainty-AIGT.

📄 PDF Abstract BibTeX arXiv:2606.02158

Code (0)

등록된 구현이 없습니다.

Tasks

Text Detection

Similar Papers 제목 키워드 기반

Attention Hijackers: Detect and Disentangle Attention Hijacking in LVLMs for Hallucination Mitigation

2025-03-11 · Beitao Chen, Xinyu Lyu, Lianli Gao, Jingkuan Song 외

Despite their success, Large Vision-Language Models (LVLMs) remain vulnerable to hallucinations. While existing studies attribute the cause of hallucinations to insufficient visual attention to image tokens, our findings…

AttributeDisentanglementHallucination

WN-Salience: A Corpus of News Articles with Entity Salience Annotations

2020-05-01 · LREC 2020 5 · Chuan Wu, Evangelos Kanoulas, Maarten de Rijke, Wei Lu

Entities can be found in various text genres, ranging from tweets and web pages to user queries submitted to web search engines. Existing research either considers all entities in the text equally important, or heuristic…

ArticlesEntity Linking

Detecting Distillation Data from Reasoning Models

2025-10-06 · Hengxiang Zhang, Hyeong Kyu Choi, Sharon Li, Hongxin Wei arxiv

Reasoning distillation has emerged as a prevailing paradigm for transferring reasoning capabilities from large reasoning models to small language models. Yet, reasoning distillation risks data contamination: benchmark da…

Table-based Fact Verification with Salience-aware Learning

2021-09-09 · Findings (EMNLP) 2021 11 · Fei Wang, Kexuan Sun, Jay Pujara, Pedro Szekely 외

Tables provide valuable knowledge that can be used to verify textual statements. While a number of works have considered table-based fact verification, direct alignments of tabular data with tokens in textual statements …

counterfactualData AugmentationFact VerificationTable-based Fact Verification

Vision-centric Token Compression in Large Language Model

2025-02-02 · Ling Xing, Alex Jinpeng Wang, Rui Yan, Xiangbo Shu 외

Real-world applications are stretching context windows to hundreds of thousand of tokens while Large Language Models (LLMs) swell from billions to trillions of parameters. This dual expansion send compute and memory cost…

In-Context LearningLanguage ModelingLanguage ModellingLarge Language Model+2