paper-with-me

홈 › Papers

Auditing Cross-Lingual Fairness in Language Model Watermarking

2026-08-20 · Alexander Nemecek, Osama Zafar, Debargha Ganguly, Vikash Singh, Vipin Chaudhary, Erman Ayday arxiv

Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English but determine conclusions cross-lingually. We propose an evaluation framework with four components: detection thresholds calibrated empirically per deployment context, a threshold-independent companion measurement that distinguishes calibration failures from detection failures, three disjoint quality measurement paradigms (distributional, paired-semantic, and reference-perplexity), and a generalized-entropy decomposition of cross-language disparity over a typological family partition. Applied to six watermarking schemes, three open-weight generators, eleven languages spanning four scripts and eight typological families, and both base and instruction-tuned regimes, the framework reveals failure modes that single-language single-paradigm evaluation cannot surface. Across detection and quality, observed disparity is predominantly between-family on the typological partition, indicating that cross-lingual fairness gaps in watermarking are structural to language properties rather than idiosyncratic to particular languages.

📄 PDF Abstract BibTeX arXiv:2608.20047

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking

2026-04-15 · Alexander Nemecek, Osama Zafar, Yuqiao Xu, Wenbiao Li 외 arxiv

Watermarking is becoming the default mechanism for AI content authentication, with governance policies and frameworks referencing it as infrastructure for content provenance. Yet across text, image, and audio modalities,…

Uncovering the Hidden Threat of Text Watermarking from Users with Cross-Lingual Knowledge

2025-02-23 · Mansour Al Ghanim, Jiaqi Xue, Rochana Prih Hastuti, Mengxin Zheng 외

In this study, we delve into the hidden threats posed to text watermarking by users with cross-lingual knowledge. While most research focuses on watermarking methods for English, there is a significant gap in evaluating …

Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to 100+ Languages via Back-Translation

2025-10-20 · Asim Mohamed, Martin Gubri arxiv

Multilingual watermarking aims to make large language model (LLM) outputs traceable across languages, yet current methods still fall short. Despite claims of cross-lingual robustness, they are evaluated only on high-reso…

Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language Models

2024-02-21 · Zhiwei He, Binglin Zhou, Hongkun Hao, Aiwei Liu 외

Text watermarking technology aims to tag and identify content produced by large language models (LLMs) to prevent misuse. In this study, we introduce the concept of cross-lingual consistency in text watermarking, which a…

TAG

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

2026-06-25 · Subham Kumar, Prakrithi Shivaprakash, Abhishek Manoharan, Astut Kurariya 외 arxiv

Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographically diverse Indian healthcare context remains largely unknown. In this study, w…

Speech Recognition