paper-with-me

Papers

Normalisation-Based Likelihood Ratio Estimation for Forensic Authorship Verification

2026-07-10 · Sadie Barlow, Andrea Nini, Edoardo Manino arxiv

Authorship verification (AV) is the task of determining whether two texts were written by the same author. In a forensic context, the strength of AV evidence can be quantified using likelihood ratios. Most AV methods are score-based and deriving well-calibrated likelihood ratios from these scores requires a separate calibration model. This, in turn, requires additional amounts of case-relevant data, which is often time-consuming to obtain and prepare. This study proposes two novel normalisation techniques, the Square Root Correction and the Hapax Correction, for deriving likelihood ratios from the AV method LambdaG without the need of a calibration model (Nini et al. 2026). These corrections are designed to mitigate the overestimation of evidential strength that may result from long or highly repetitive texts. Performance is evaluated against logistic regression calibration across fifteen corpora and a range of text lengths (100-9,500 tokens), using the log-likelihood ratio cost (Cllr). The proposed methods achieve performance comparable to logistic regression calibration, with the Hapax Correction outperforming it in approximately 45% of tests (weighted by corpora). Furthermore, performance was more frequently close (within 5%) when the Hapax Correction was outperformed by logistic regression calibration, compared with the reverse comparison. Eliminating the need to train a calibration model reduces data-requirements, time and complexity, thereby increasing the accessibility and transparency of forensic text comparison. This combination of empirical performance and practical advantages supports the adoption of the proposed methods in forensic settings.

📄 PDF Abstract BibTeX arXiv:2607.09501

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Feature-Based Forensic Text Comparison Using a Poisson Model for Likelihood Ratio Estimation

2020-12-01 · ALTA 2020 12 · Michael Carne, Shunichi Ishihara

Score- and feature-based methods are the two main ones for estimating a forensic likelihood ratio (LR) quantifying the strength of evidence. In this forensic text comparison (FTC) study, a score-based method using the Co…

Authorship Attributionfeature selection

Fusing Stylometric and Embedding Systems to Estimate Authorship Likelihood Ratios in Japanese

2026-06-12 · Praju Ghatpande, Satoru Tsuge, Shunichi Ishihara, Wataru Zaitsu 외 arxiv

The likelihood ratio framework is widely recognized as the logically and legally sound basis for evidential analysis across forensic sciences, and its importance is increasingly acknowledged in analyses of authorship in …

Authorship Impersonation via LLM Prompting does not Evade Authorship Verification Methods

2026-03-31 · Baoyi Zeng, Andrea Nini arxiv

Authorship verification (AV), the task of determining whether a questioned text was written by a specific individual, is a critical part of forensic linguistics. While manual authorial impersonation by perpetrators has l…

Learning Energy-Based Models by Self-normalising the Likelihood

2025-03-10 · Hugo Senetaire, Paul Jeha, Pierre-Alexandre Mattei, Jes Frellsen

Training an energy-based model (EBM) with maximum likelihood is challenging due to the intractable normalisation constant. Traditional methods rely on expensive Markov chain Monte Carlo (MCMC) sampling to estimate the gr…

Density Estimation

WEKA in Forensic Authorship Analysis: A corpus-based approach of Saudi Authors

2020-12-01 · ICON 2020 12 · Mashael AlAmr, Eric Atwell

This is a pilot study that aims to explore the potential of using WEKA in forensic authorship analysis. It is a corpus-based research using data from Twitter collected from thirteen authors from Riyadh, Saudi Arabia. It …