paper-with-me

홈 › Papers

Ideal Attribution and Faithful Watermarks for Language Models

2025-12-07 · Min Jae Song, Kameron Shahabi arxiv

We introduce ideal attribution mechanisms, a formal abstraction for reasoning about attribution decisions over strings. At the core of this abstraction lies the ledger, an append-only log of the prompt-response interaction history between a model and its user. Each mechanism produces deterministic decisions based on the ledger and an explicit selection criterion, making it well-suited to serve as a ground truth for attribution. We frame the design goal of watermarking schemes as faithful representation of ideal attribution mechanisms. This novel perspective brings conceptual clarity, replacing piecemeal probabilistic statements with a unified language for stating the guarantees of each scheme. It also enables precise reasoning about desiderata for future watermarking schemes, even when no current construction achieves them, since the ideal functionalities are specified first. In this way, the framework provides a roadmap that clarifies which guarantees are attainable in an idealized setting and worth pursuing in practice.

📄 PDF Abstract BibTeX arXiv:2512.07038

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Watermark-based Attribution of AI-Generated Content

2024-04-05 · Zhengyuan Jiang, Moyang Guo, Yuepeng Hu, Neil Zhenqiang Gong

Several companies have deployed watermark-based detection to identify AI-generated content. However, attribution--the ability to trace back to the user of a generative AI (GenAI) service who created a given piece of AI-g…

A Multilingual Perspective Towards the Evaluation of Attribution Methods in Natural Language Inference

2022-04-11 · Kerem Zaman, Yonatan Belinkov

Most evaluations of attribution methods focus on the English language. In this work, we present a multilingual approach for evaluating attribution methods for the Natural Language Inference (NLI) task in terms of faithfu…

Natural Language Inference

ProMark: Proactive Diffusion Watermarking for Causal Attribution

2024-03-14 · CVPR 2024 1 · Vishal Asnani, John Collomosse, Tu Bui, Xiaoming Liu 외

Generative AI (GenAI) is transforming creative workflows through the capability to synthesize and manipulate images via high-level prompts. Yet creatives are not well supported to receive recognition or reward for the us…

Attribute

A Multilingual Perspective Towards the Evaluation of Attribution Methods

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Most evaluations of attribution methods focus on the English language. In this work, we present a multilingual approach for evaluating attribution methods for the Natural Language Inference (NLI) task in terms of plausib…

Natural Language Inference

RSAT: Structured Attribution Makes Small Language Models Faithful Table Reasoners

2026-04-30 · Jugal Gajjar, Kamalasankari Subramaniakuppusamy arxiv

When a language model answers a table question, users have no way to verify which cells informed which reasoning steps. We introduce RSAT, a method that trains small language models (SLMs, 1-8B) to produce step-by-step r…