paper-with-me

Papers Model extraction

“Model extraction” 태그가 달린 논문 218편 · 필터 해제

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

2026-08-27 · Chen Chen, Yaolin Chen, Xuehan Sun, Juan Lin 외 arxiv

Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to mode…

Model extraction

Caliber: Cross-Architecture Extraction-Cost Control for Score-Returning APIs

2026-08-02 · Chi Wang, Hanwen Wang, Yu Xia, Zihan Wang 외 arxiv

We present Caliber, an output-perturbation defense against model extraction that formulates noise selection as a calibration problem: how much the defense degrades the supervision signal used to train a surrogate, and th…

Knowledge DistillationModel extraction

DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods

2026-07-17 · Jens Frankenreiter arxiv

Much empirical legal research depends on translating unstructured text into structured variables. In corporate governance research as elsewhere, this translation has traditionally relied on human coding of documents such…

Binary ClassificationModel extraction

Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot

2026-06-14 · Yuyang Dai, Yushun Dong arxiv

Large language models deployed as commercial APIs are vulnerable to model extraction attacks, while existing defenses either act too late or degrade utility for legitimate users. We propose \textbf{Knowledge Trap}, a def…

Model extraction

SciR: A Controllable Benchmark for Scientific Reasoning in LLMs

2026-06-11 · Pierre Beckmann, Marco Valentino, Andre Freitas arxiv

Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is currently out of reach: scientific benchmar…

Model extraction

T2S: A Rehearsal-Based Approach for Extraction-Resistant Model Watermarking

2026-06-10 · Jian-Ping Mei, Weibin Zhang, Ao Yao, Tiantian Zhu 외 arxiv

Model watermarking safeguards AI model intellectual property by embedding distinctive knowledge that induces unique behavioral signatures. The primary technical challenge lies in ensuring watermark robustness against var…

Model extraction

An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic

2026-06-04 · Shuze Liu, Qianwen Guo, Yushun Dong arxiv

Large language models (LLMs) are increasingly deployed through hosted APIs, making model extraction a practical threat to model ownership and service security. However, individual extraction queries often resemble benign…

Model extraction

BAHSD: Bridging the Long-tail Gap via Adaptive Distillation in Black-box Sequential Recommendation

2026-06-02 · Xi Zhou, Famin Wu, Mingming Li, Hongyue Zhang 외 arxiv

Sequential recommendation systems are widely adopted but often deployed as black-box APIs, which has driven recent interest in model extraction to replicate their capabilities locally. However, the long-tail distribution…

Sequential RecommendationContrastive LearningModel extraction

AI Model Extraction Attacks: Bypassing Single-Client Assumptions in Defenses

2026-06-02 · Maxime Schwarzer, Johannes F. Loevenich, Gustavo Sánchez, Laurin Holz 외 arxiv

Ensuring the protection of Artificial Intelligence (AI) models deployed in military Command and Control (C2) systems and critical infrastructure is essential for maintaining information superiority. Model Extraction Atta…

Model extraction

FlowGuard: Flow Matching for Identity-Independent Detection of Data-Free Model Stealing Attacks on Energy System Intrusion Detection Systems

2026-06-02 · Maxime Schwarzer, Laurin Holz, Tobias Huerten, Johannes Loevenich 외 arxiv

Artificial Intelligence (AI)-based Intrusion Detection Systems (IDS) deployed in energy infrastructure are vulnerable to model theft attacks, which allow adversaries to create evasive traffic offline. Current defences ag…

Intrusion DetectionModel extraction

A Registry-Bound LLM Pipeline for Evidence-Grounded Trait Extraction across Tropical Plants, Aquatic Species, and Exotic Pets

2026-05-31 · Jeff Wang arxiv

We describe a registry-bound large-language-model extraction pipeline producing evidence-grounded structured trait records at scale, on cultivated tropical plant, aquatic, and pet species. Four mechanisms render LLM-deri…

Model extraction

Can Subgraph Explanations Be Weaponized to Steal Graph Neural Networks?

2026-05-28 · Ojas Nimase, Jiate Li, Yue Zhao, Yushun Dong arxiv

Graph Machine Learning as a Service (GMLaaS) platforms increasingly implement explainability interfaces to meet regulatory transparency requirements. However, this transparency creates exploitable vulnerabilities for mod…

Graph ClassificationModel extraction

EMMA: Extracting Multiple physical parameters from Multimodal Data

2026-05-21 · Farhat Shaikh, Ayan Banerjee, Sandeep Gupta arxiv

We introduce EMMA, a physics-informed multimodal framework that recovers all identifiable dynamical parameters of a system directly from raw video, audio, and image-based time-series observations. Unlike prior video-only…

Model extraction

Adaptive Probe-based Steering for Robust LLM Jailbreaking

2026-05-19 · Junxi Chen, Junhao Dong, Xiaohua Xie arxiv

Recent work has demonstrated the potential of contrastive steering for jailbreaking Large Language Models (LLMs). However, existing methods rely on limited and inherently biased contrastive prompts and require laborious …

Model extraction

MADP: A Multi-Agent Pipeline for Sustainable Document Processing with Human-in-the-Loop

2026-05-16 · Diego Gosmar, Giovanni Zenezini arxiv

Document processing automation remains a critical challenge in enterprise environments, where traditional manual approaches are labor-intensive and error-prone. We present MADP, a multi-agent architecture that addresses …

Model extraction

Identified-Set Geometry of Distributional Model Extraction under Top-$K$ Censored API Access

2026-05-11 · Wenhua Nie, ZiCheng Zhu, Jianan Wu, Binhan Luo 외 arxiv

Modern LLM APIs often reveal only top-$K$ logit scores and censor the remaining vocabulary. We study the per-position distribution-recovery limits of this access model. For censoring threshold $τ$, the compatible teacher…

Model extraction

TrEEStealer: Stealing Decision Trees via Enclave Side Channels

2026-04-20 · Jonas Sander, Anja Rabich, Nick Mahling, Felix Maurer 외 arxiv

Today, machine learning is widely applied in sensitive, security-related, and financially lucrative applications. Model extraction attacks undermine current business models where a model owner sells model access, e.g., v…

Model extraction

A Public Theory of Distillation Resistance via Constraint-Coupled Reasoning Architectures

2026-03-26 · Peng Wei, Wesley Shu arxiv

Knowledge distillation, model extraction, and behavior transfer have become central concerns in frontier AI. The main risk is not merely copying, but the possibility that useful capability can be transferred more cheaply…

Knowledge DistillationModel extraction

AI Security in the Foundation Model Era: A Comprehensive Survey from a Unified Perspective

2026-03-25 · Zhenyi Wang, Siyu Luan arxiv

As machine learning (ML) systems expand in both scale and functionality, the security landscape has become increasingly complex, with a proliferation of attacks and defenses. However, existing studies largely treat these…

Model extraction

TabKD: Tabular Knowledge Distillation through Interaction Diversity of Learned Feature Bins

2026-03-16 · Shovon Niverd Pereira, Krishna Khadka, Yu Lei arxiv

Data-free knowledge distillation enables model compression without original training data, critical for privacy-sensitive tabular domains. However, existing methods does not perform well on tabular data because they do n…

Data-free Knowledge DistillationModel CompressionModel extraction
1–20 / 218 다음 →