paper-with-me

홈 › Papers

KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing

2026-05-28 · Yijia Fang, Yiqing Feng, Bingyu Li, Mingxun Zhou arxiv

Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a claimed endpoint is actually serving the advertised model. We introduce KBF, a low-cost black-box auditing protocol that fingerprints model APIs using stable numerical recall near the knowledge boundary. Across 16 production LLM endpoints, KBF flags all 155 economically relevant substitutions without rejecting any same-model controls, remains stable under deployment variation, detects high-separation mixed-routing attacks when only 5-10% of traffic is substituted, and finds that 7 of 27 platform model cells in a six-platform shadow API audit are statistically inconsistent with their reference endpoints, with inconsistencies concentrated on premium Claude endpoints.

📄 PDF Abstract BibTeX arXiv:2605.29524

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SoK: Large Language Model Copyright Auditing via Fingerprinting

2025-08-27 · Shuo Shao, Yiming Li, Yu He, Hongwei Yao 외 arxiv

The broad capabilities and substantial resources required to train Large Language Models (LLMs) make them valuable intellectual property, yet they remain vulnerable to copyright infringement, such as unauthorized use and…

SDBF: Steep-Decision-Boundary Fingerprinting for Hard-Label Tampering Detection of DNN Models

2025-01-01 · CVPR 2025 1 · Xiaofan Bai, Shixin Li, Xiaojing Ma, Bin Benjamin Zhu 외

Cloud-based AI systems offer significant benefits but also introduce vulnerabilities, making deep neural network (DNN) models susceptible to malicious tampering. This tampering may involve harmful behavior injection …

Sensitivity

CALM: Curiosity-Driven Auditing for Large Language Models

2025-01-06 · Xiang Zheng, Longxiang Wang, Yi Liu, Xingjun Ma 외

Auditing Large Language Models (LLMs) is a crucial and challenging task. In this study, we focus on auditing black-box LLMs without access to their parameters, only to the provided service. We treat this type of auditing…

PreUnlearn: Auditing Collateral Knowledge Damage Before Large Language Model Unlearning

2026-06-16 · Bo Su, Ankit Shah, Thai Le arxiv

Machine unlearning for large language models (LLMs) aims to remove specified knowledge while preserving the rest of the model's capabilities. However, the boundary between knowledge to forget and knowledge to retain is o…

EditMF: Drawing an Invisible Fingerprint for Your Large Language Models

2025-08-12 · Jiaxuan Wu, Yinghan Zhou, Wanli Peng, Yiming Xue 외 arxiv

Training large language models (LLMs) is resource-intensive and expensive, making protecting intellectual property (IP) for LLMs crucial. Recently, embedding fingerprints into LLMs has emerged as a prevalent method for e…