paper-with-me

홈 › Papers

Distill-and-Compare: Auditing Black-Box Models Using Transparent Model Distillation

2017-10-17 · Sarah Tan, Rich Caruana, Giles Hooker, Yin Lou

Black-box risk scoring models permeate our lives, yet are typically proprietary or opaque. We propose Distill-and-Compare, a model distillation and comparison approach to audit such models. To gain insight into black-box models, we treat them as teachers, training transparent student models to mimic the risk scores assigned by black-box models. We compare the student model trained with distillation to a second un-distilled transparent model trained on ground-truth outcomes, and use differences between the two models to gain insight into the black-box model. Our approach can be applied in a realistic setting, without probing the black-box model API. We demonstrate the approach on four public data sets: COMPAS, Stop-and-Frisk, Chicago Police, and Lending Club. We also propose a statistical test to determine if a data set is missing key features used to train the black-box model. Our test finds that the ProPublica data is likely missing key feature(s) used in COMPAS.

📄 PDF Abstract BibTeX arXiv:1710.06169

Code (1)

shftan/auditblackbox

Similar Papers 제목 키워드 기반

TRUST: A Decentralized Framework for Auditing Large Language Model Reasoning

2025-10-23 · Morris Yu-Chao Huang, Zhen Tan, Mohan Zhang, Pingzhi Li 외 arxiv

Large Language Models generate complex reasoning chains that reveal their decision-making, yet verifying the faithfulness and harmlessness of these intermediate steps remains a critical unsolved problem. Existing auditin…

Black-Box On-Policy Distillation of Large Language Models

2025-11-13 · Tianzhu Ye, Li Dong, Zewen Chi, Xun Wu 외 arxiv

Black-box distillation creates student large language models (LLMs) by learning from a proprietary teacher model's text outputs alone, without access to its internal logits or parameters. In this work, we introduce Gener…

Knowledge Distillation

ILDiff: Generate Transparent Animated Stickers by Implicit Layout Distillation

2024-12-30 · Ting Zhang, Zhiqiang Yuan, Yeshuang Zhu, Jinchao Zhang

High-quality animated stickers usually contain transparent channels, which are often ignored by current video generation models. To generate fine-grained animated transparency channels, existing methods can be roughly di…

Image MattingVideo GenerationVideo Matting

Distillation Quantification for Large Language Models

2025-01-22 · Sunbowen Lee, Junting Zhou, Chang Ao, Kaige Li 외

Model distillation is a technique for transferring knowledge from large language models (LLMs) to smaller ones, aiming to create resource-efficient yet high-performing models. However, excessive distillation can lead to …

Training Domain Draft Models for Speculative Decoding: Best Practices and Insights

2025-03-10 · Fenglu Hong, Ravi Raju, Jonathan Lingjie Li, Bo Li 외

Speculative decoding is an effective method for accelerating inference of large language models (LLMs) by employing a small draft model to predict the output of a target model. However, when adapting speculative decoding…

Knowledge Distillation