paper-with-me

Papers

Machine Learning for Detecting Data Exfiltration: A Review

2020-12-17 · Bushra Sabir, Faheem Ullah, M. Ali Babar, Raj Gaire

Context: Research at the intersection of cybersecurity, Machine Learning (ML), and Software Engineering (SE) has recently taken significant steps in proposing countermeasures for detecting sophisticated data exfiltration attacks. It is important to systematically review and synthesize the ML-based data exfiltration countermeasures for building a body of knowledge on this important topic. Objective: This paper aims at systematically reviewing ML-based data exfiltration countermeasures to identify and classify ML approaches, feature engineering techniques, evaluation datasets, and performance metrics used for these countermeasures. This review also aims at identifying gaps in research on ML-based data exfiltration countermeasures. Method: We used a Systematic Literature Review (SLR) method to select and review {92} papers. Results: The review has enabled us to (a) classify the ML approaches used in the countermeasures into data-driven, and behaviour-driven approaches, (b) categorize features into six types: behavioural, content-based, statistical, syntactical, spatial and temporal, (c) classify the evaluation datasets into simulated, synthesized, and real datasets and (d) identify 11 performance measures used by these studies. Conclusion: We conclude that: (i) the integration of data-driven and behaviour-driven approaches should be explored; (ii) There is a need of developing high quality and large size evaluation datasets; (iii) Incremental ML model training should be incorporated in countermeasures; (iv) resilience to adversarial learning should be considered and explored during the development of countermeasures to avoid poisoning attacks; and (v) the use of automated feature engineering should be encouraged for efficiently detecting data exfiltration attacks.

📄 PDF Abstract BibTeX arXiv:2012.09344

Code (0)

등록된 구현이 없습니다.

Tasks

Automated Feature EngineeringBIG-bench Machine LearningFeature EngineeringSystematic Literature Review

Similar Papers 제목 키워드 기반

Reframing Threat Detection: Inside esINSIDER

2019-04-07 · M. Arthur Munson, Jason Kichen, Dustin Hillard, Ashley Fidler 외

We describe the motivation and design for esINSIDER, an automated tool that detects potential persistent and insider threats in a network. esINSIDER aggregates clues from log data, over extended time periods, and propose…

BIG-bench Machine Learning

Evasion-Resilient Detection of DNS-over-HTTPS Data Exfiltration: A Practical Evaluation and Toolkit

2025-12-23 · Adam Elaoumari arxiv

The purpose of this project is to assess how well defenders can detect DNS-over-HTTPS (DoH) file exfiltration, and which evasion strategies can be used by attackers. While providing a reproducible toolkit to generate, in…

Unsupervised Learning of Distributional Properties can Supplement Human Labeling and Increase Active Learning Efficiency in Anomaly Detection

2023-07-13 · Jaturong Kongmanee, Mark Chignell, Khilan Jerath, Abhay Raman

Exfiltration of data via email is a serious cybersecurity threat for many organizations. Detecting data exfiltration (anomaly) patterns typically requires labeling, most often done by a human annotator, to reduce the hig…

Active LearningAnomaly DetectionUnsupervised Anomaly Detection

Aggressive Compression Enables LLM Weight Theft

2026-01-03 · Davis Brown, Juan-Pablo Rivera, Dan Hendrycks, Mantas Mazeika arxiv

As frontier AIs become more powerful and costly to develop, adversaries have increasing incentives to steal model weights by mounting exfiltration attacks. In this work, we consider exfiltration attacks where an adversar…

Your LLM Agent Can Leak Your Data: Data Exfiltration via Backdoored Tool Use

2026-04-07 · Wuyang Zhang, Shichao Pei arxiv

Tool-use large language model (LLM) agents are increasingly deployed to support sensitive workflows, relying on tool calls for retrieval, external API access, and session memory management. While prior research has exami…