paper-with-me

Papers

Verifying Classification with Limited Disclosure

2025-02-22 · Siddharth Bhandari, Liren Shan

We consider the multi-party classification problem introduced by Dong, Hartline, and Vijayaraghavan (2022) motivated by electronic discovery. In this problem, our goal is to design a protocol that guarantees the requesting party receives nearly all responsive documents while minimizing the disclosure of nonresponsive documents. We develop verification protocols that certify the correctness of a classifier by disclosing a few nonresponsive documents. We introduce a combinatorial notion called the Leave-One-Out dimension of a family of classifiers and show that the number of nonresponsive documents disclosed by our protocol is at most this dimension in the realizable setting, where a perfect classifier exists in this family. For linear classifiers with a margin, we characterize the trade-off between the margin and the number of nonresponsive documents that must be disclosed for verification. Specifically, we establish a trichotomy in this requirement: for $d$ dimensional instances, when the margin exceeds $1/3$, verification can be achieved by revealing only $O(1)$ nonresponsive documents; when the margin is exactly $1/3$, in the worst case, at least $\Omega(d)$ nonresponsive documents must be disclosed; when the margin is smaller than $1/3$, verification requires $\Omega(e^d)$ nonresponsive documents. We believe this result is of independent interest with applications to coding theory and combinatorial geometry. We further extend our protocols to the nonrealizable setting defining an analogous combinatorial quantity robust Leave-One-Out dimension, and to scenarios where the protocol is tolerant to misclassification errors by Alice.

📄 PDF Abstract BibTeX arXiv:2502.16352

Code (0)

등록된 구현이 없습니다.

Tasks

Classification

Similar Papers 제목 키워드 기반

Selective Disclosure Watermarking for Large Language Models

2026-07-06 · Xuyang Chen, Xiang Li, Yangxinyu Xie, Qi Long arxiv

Watermarking methods embed imperceptible and verifiable signals into text generated by large language models (LLMs). Existing approaches include zero-bit schemes for distinguishing synthetic text from human writing and m…

Identifying Medical Self-Disclosure in Online Communities

2021-06-01 · NAACL 2021 4 · Mina Valizadeh, Pardis Ranjbar-Noiey, Cornelia Caragea, Natalie Parde

Self-disclosure in online health conversations may offer a host of benefits, including earlier detection and treatment of medical issues that may have otherwise gone unaddressed. However, research analyzing medical self-…

Evaluating TCFD Reporting: A New Application of Zero-Shot Analysis to Climate-Related Financial Disclosures

2023-02-01 · Alix Auzepy, Elena Tönjes, David Lenz, Christoph Funk

We examine climate-related disclosures in a large sample of reports published by banks that officially endorsed the recommendations of the Task Force for Climate-related Financial Disclosures (TCFD). In doing so, we intr…

text-classificationText ClassificationZero-Shot Text Classification

A Semantics-based Approach to Disclosure Classification in User-Generated Online Content

2020-11-01 · Findings of the Association for Computational Linguistics 2020 · Chandan Akiti, Anna Squicciarini, Sarah Rajtmajer

As users engage in public discourse, the rate of voluntarily disclosed personal information has seen a steep increase. So-called self-disclosure can result in a number of privacy concerns. Users are often unaware of the …

Semantic Role Labelingtext-classificationText Classification

Assessing Crime Disclosure Patterns in a Large-Scale Cybercrime Forum

2026-03-02 · Raphael Hoheisel, Tom Meurs, Jai Wientjes, Marianne Junger 외 arxiv

Cybercrime forums play a central role in the cybercrime ecosystem, serving as hubs for the exchange of illicit goods, services, and knowledge. Previous studies have explored the market and social structures of these foru…

Text Classification