The Sweet Danger of Sugar: Debunking Representation Learning for Encrypted Traffic Classification
Recently we have witnessed the explosion of proposals that, inspired by Language Models like BERT, exploit Representation Learning models to create traffic representations. All of them promise astonishing performance in encrypted traffic classification (up to 98% accuracy). In this paper, with a networking expert mindset, we critically reassess their performance. Through extensive analysis, we demonstrate that the reported successes are heavily influenced by data preparation problems, which allow these models to find easy shortcuts - spurious correlation between features and labels - during fine-tuning that unrealistically boost their performance. When such shortcuts are not present - as in real scenarios - these models perform poorly. We also introduce Pcap-Encoder, an LM-based representation learning model that we specifically design to extract features from protocol headers. Pcap-Encoder appears to be the only model that provides an instrumental representation for traffic classification. Yet, its complexity questions its applicability in practical settings. Our findings reveal flaws in dataset preparation and model training, calling for a better and more conscious test design. We propose a correct evaluation methodology and stress the need for rigorous benchmarking.
Code (0)
등록된 구현이 없습니다.
Tasks
Representation LearningSimilar Papers 제목 키워드 기반
The efficacy of the sugar-free labels is reduced by the health-sweetness tradeoff
In the present study, we use an experimental setting to explore the effects of sugar-free labels on the willingness to pay for food products. In our experiment, participants placed bids for sugar-containing and analogous…
Prenatal Sugar Consumption and Late-Life Human Capital and Health: Analyses Based on Postwar Rationing and Polygenic Scores
Maternal sugar consumption in utero may have a variety of effects on offspring. We exploit the abolishment of the rationing of sweet confectionery in the UK on April 24, 1949, and its subsequent reintroduction some month…
SUGAR: A Sweeter Spot for Generative Unlearning of Many Identities
Recent advances in 3D-aware generative models have enabled high-fidelity image synthesis of human identities. However, this progress raises urgent questions around user consent and the ability to remove specific individu…
Accelerating Experimental Design by Incorporating Experimenter Hunches
Experimental design is a process of obtaining a product with target property via experimentation. Bayesian optimization offers a sample-efficient tool for experimental design when experiments are expensive. Often, expert…
Bayesian OptimizationExperimental DesignSplash! Robustifying Donor Pools for Policy Studies
Policy researchers using synthetic control methods typically choose a donor pool in part by using policy domain expertise so the untreated units are most like the treated unit in the pre intervention period. This potenti…