The Thieves on Sesame Street are Polyglots - Extracting Multilingual Models from Monolingual APIs
Pre-training in natural language processing makes it easier for an adversary with only query access to a victim model to reconstruct a local copy of the victim by training with gibberish input data paired with the victim{'}s labels for that data. We discover that this extraction process extends to local copies initialized from a pre-trained, multilingual model while the victim remains monolingual. The extracted model learns the task from the monolingual victim, but it generalizes far better than the victim to several other languages. This is done without ever showing the multilingual, extracted model a well-formed input in any of the languages for the target task. We also demonstrate that a few real examples can greatly improve performance, and we analyze how these results shed light on how such extraction methods succeed.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Code-Mixing on Sesame Street: Dawn of the Adversarial Polyglots
Multilingual models have demonstrated impressive cross-lingual transfer performance. However, test sets like XNLI are monolingual at the example level. In multilingual communities, it is common for polyglots to code-mix …
Cross-Lingual TransferXLM-RThieves on Sesame Street! Model Extraction of BERT-based APIs
We study the problem of model extraction in natural language processing, in which an adversary with only query access to a victim model attempts to reconstruct a local copy of that model. Assuming that both the adversary…
Language ModelingLanguage ModellingModel extractionNatural Language Inference+2Sesame Street to Mount Sinai: BERT-constrained character-level Moses models for multilingual lexical normalization
This paper describes the HEL-LJU submissions to the MultiLexNorm shared task on multilingual lexical normalization. Our system is based on a BERT token classification preprocessing step, where for each token the type of …
Lexical Normalizationtoken-classificationToken ClassificationSingle Headed Attention RNN: Stop Thinking With Your Head
The leading approaches in language modeling are all obsessed with TV shows of my youth - namely Transformers and Sesame Street. Transformers this, Transformers that, and over here a bonfire worth of GPU-TPU-neuromorphic …
GPUHyperparameter OptimizationLanguage ModelingLanguage ModellingGREEK-BERT: The Greeks visiting Sesame Street
Transformer-based language models, such as BERT and its variants, have achieved state-of-the-art performance in several downstream natural language processing (NLP) tasks on generic benchmark datasets (e.g., GLUE, SQUAD,…
Language ModelingLanguage Modellingnamed-entity-recognitionNamed Entity Recognition+5