Defending computer networks from cyber threats becomes more difficult each year as their quantity, complexity, and scope increases constantly. These networks cannot be protected by human cyber security experts alone. Machine learning tools, such as intrusion detection and prevention systems, can be deployed everywhere, are easily updated, and never sleep. The research necessary for greater adoption of these tools, however, is held back by datasets that are sparse, and often unlabeled and unbalanced. Generative models present a solution to these problems, but they present further research questions, namely:
- How do we adapt generative models to the myriad of highly categorical feature representations of cyber security data?- Which types of generative models will produce the highest quality cyber security data?
- How do we evaluate the quality of this data?
- How do we determine which types of data will be most relevant in developing an effective cyber defense?
The objective of this PhD research is to answer each of these research questions and accelerate the further development of machine learning applications for cyber security. It accomplishes this using five contributions. The first of these is a survey and synthesis of generative models applied to intrusion detection and related tasks. The survey, through both a systematic mapping study and literature review, offers insights into the current state of generative models around these questions and provides directions for further research.
The second contribution, inspired by these directions, is an evaluation of denoising diffusion implicit models for generating cyber security data, alongside an evaluation of different feature representations used with generative models. This is further expanded upon in the third contribution, which demonstrates the use of conformal prediction to confidently select optimal generative models for producing cyber security data on a per-class basis.
The fourth contribution of this research focuses on the question of data relevance. It provides a suite of metrics that are designed to measure the ability of a network's sensors to detect an attack. These metrics feature a degree of modularity, both to make them more compatible with a variety of different types of networks, but also to render them future-proof as newer innovations in attacker threat modeling emerge.
The final contribution provides a method for solving the higher level goal of data availability without the use of generative models. This is presented as an algorithm for generating synthetic network topologies for feeder networks. As there is little to no training data for descriptions of feeder network topologies, a customized generator is required.
Future cyber defenses will depend upon data that is balanced, labeled, relevant, and of a high quality. This research provides the metrics, algorithms, and applications that pave the way to this reality.
Metrics
1 Record Views
Details
Title
Machine Learning for Network Security: Metrics, Algorithms, and Applications
Creators
James Michael Melim Halvorsen
Contributors
Assefaw H Gebremedhin (Advisor)
Janardhan Doppa (Committee Member)
Monowar Hasan (Committee Member)
Awarding Institution
Washington State University
Academic Unit
School of Electrical Engineering and Computer Science
Theses and Dissertations
Doctor of Philosophy (PhD), Washington State University