paper-with-me

홈 › Papers

DMAP: A Distribution Map for Text

2026-02-12 · Tom Kempton, Julia Rozanova, Parameswaran Kamalaruban, Maeve Madigan, Karolina Wresilo, Yoann L. Launay, David Sutton, Stuart Burrell arxiv

Large Language Models (LLMs) are a powerful tool for statistical text analysis, with derived sequences of next-token probability distributions offering a wealth of information. Extracting this signal typically relies on metrics such as perplexity, which do not adequately account for context; how one should interpret a given next-token probability is dependent on the number of reasonable choices encoded by the shape of the conditional distribution. In this work, we present DMAP, a mathematically grounded method that maps a text, via a language model, to a set of samples in the unit interval that jointly encode rank and probability information. This representation enables efficient, model-agnostic analysis and supports a range of applications. We illustrate its utility through three case studies: (i) validation of generation parameters to ensure data integrity, (ii) examining the role of probability curvature in machine-generated text detection, and (iii) a forensic analysis revealing statistical fingerprints left in downstream models that have been subject to post-training on synthetic data. Our results demonstrate that DMAP offers a unified statistical view of text that is simple to compute on consumer hardware, widely applicable, and provides a foundation for further research into text analysis with LLMs.

📄 PDF Abstract BibTeX arXiv:2602.11871

Code (0)

등록된 구현이 없습니다.

Tasks

Text Detection

Similar Papers 제목 키워드 기반

Text to Multi-level MindMaps: A Novel Method for Hierarchical Visual Abstraction of Natural Language Text

2014-08-01 · Mohamed Elhoseiny, Ahmed Elgammal

MindMapping is a well-known technique used in note taking, which encourages learning and studying. MindMapping has been manually adopted to help present knowledge and concepts in a visual form. Unfortunately, there is no…

3D Modality-Aware Pre-training for Vision-Language Model in MRI Multi-organ Abnormality Detection

2026-02-27 · Haowen Zhu, Ning Yin, Xiaogen Zhou arxiv

Vision-language models (VLMs) show strong potential for complex diagnostic tasks in medical imaging. However, applying VLMs to multi-organ medical imaging introduces two principal challenges: (1) modality-specific vision…

Representation Learning

P-MapNet: Far-seeing Map Generator Enhanced by both SDMap and HDMap Priors

2024-03-15 · Zhou Jiang, Zhenxin Zhu, Pengfei Li, Huan-ang Gao 외

Autonomous vehicles are gradually entering city roads today, with the help of high-definition maps (HDMaps). However, the reliance on HDMaps prevents autonomous vehicles from stepping into regions without this expensive …

Autonomous Vehicles

GradMAP: Gradient-Based Multi-Agent Proximal Learning for Grid-Edge Flexibility

2026-04-27 · Yihong Zhou, Hongtai Zeng, Thomas Morstyn arxiv

Coordinating large populations of grid-edge devices requires learning methods that remain fully decentralised in deployment while still respecting three-phase AC distribution-network physics. This paper proposes gradient…

Self-Supervised Learning

PrevPredMap: Exploring Temporal Modeling with Previous Predictions for Online Vectorized HD Map Construction

2024-07-24 · Nan Peng, Xun Zhou, Mingming Wang, Xiaojun Yang 외

Temporal information is crucial for detecting occluded instances. Existing temporal representations have progressed from BEV or PV features to more compact query features. Compared to these aforementioned features, predi…

DecoderOnline Vectorized HD Map ConstructionPosition