PAIRWISE CLUSTERS OPTIMIZATION AND CLUSTER MOST SIGNIFICANT FEATURE METHODS FOR ANOMALY-BASED NETWORK INTRUSION DETECTION SYSTEM (POC2MSF)

Authors

  • Gervais Hatungimana Department of Informatics Engineering, Institut Teknologi Sepuluh Nopember, Indonesia

DOI:

https://doi.org/10.24191/mjoc.v3i2.3598

Keywords:

Clustering, Cluster Most Significant Feature, Network Traffic Baseline, Network Security, Quality Threshold

Abstract

Anomaly-based Intrusion Detection System (IDS) uses known baseline to detect patterns which have deviated from normal behavior. If the baseline is faulty, the IDS performance degrades. Most of researches in IDS which use k-centroids-based clustering methods like K-means, K-medoids, Fuzzy, Hierarchical and agglomerative algorithms to baseline network traffic suffer from high false positive rate compared to signature-based IDS, simply because the nature of these algorithms risk to force some network traffic into wrong profiles depending on K number of clusters needed. In this paper, we propose an alternative method which instead of defining K number of clusters, defines t distance threshold. The unrecognizable IDS; IDS which is neither HIDS nor NIDS is the consequence of using statistical methods for features selection. The speed, memory and accuracy of IDS are affected by inappropriate features reduction method or ignorance of irrelevant features. In this paper, we use two-step features selection and Quality Threshold with Optimization methods to design anomaly-based HIDS and NIDS separately. The performance of our system is 0% ,99.99%, 1,1 false positive rates, accuracy, precision and recall respectively for NIDS and 0%,99.61%, 0.991,0.97 false positive rates, accuracy, precision and recall respectively for HIDS.

References

Aggarwal, P., & Kumar, S. (2015). Analysis of KDD Dataset Attributes - Class wise For Intrusion Detection. Procedia - Procedia Computer Science, 57, 842–851. https://doi.org/10.1016/j.procs.2015.07.490

Agrawal, S., & Agrawal, J. (2015). Survey on Anomaly Detection using Data Mining Techniques. Procedia - Procedia Computer Science, 60, 708–713. https://doi.org/10.1016/j.procs.2015.08.220

Al-Mamory, S. O., & Jassim, F. S. (2015). On the designing of two grains levels network intrusion detection system. Karbala International Journal of Modern Science, 1(1), 15–

25. https://doi.org/10.1016/j.kijoms.2015.07.002

Chitrakar, R., & Chuanhe, H. (2012). Anomaly Detection using Support Vector Machine Classification with k-Medoids Clustering. In 4th International Conference on Computing and Informatics, ICOCI, Sarawak, Malaysia (pp. 1–5). https://doi.org/10.1109/AHICI.2012.6408446

Davidoff, S, Jonathan, H. (2012). Network Forensics. Upper Saddle River, New Jersey: PRENTICE HALL.

Davis, J. J., & Clark, A. J. (2011). Data preprocessing for anomaly based network intrusion detection: A review. Computers {&} Security, 30(6–7), 353–375. https://doi.org/10.1016/j.cose.2011.05.008

Fossaceca, J. M., Mazzuchi, T. A., & Sarkani, S. (2015). MARK-ELM : Application of a novel Multiple Kernel Learning framework for improving the robustness of Network Intrusion Detection. Expert Systems With Applications, 42(8), 4062–4080. https://doi.org/10.1016/j.eswa.2014.12.040

Fries, T. P. (2015). Fuzzy Clustering of Network Traffic Features for Security. IEEE Symposium on Large Data Analysis and Visualization 2015, 127–128.

Gervais, H., Munif, A., & Ahmad, T. (2016). Using QualityThreshold Distance to Detect Intrusion in TCP / IP Network. In 2016 IEEE International Conference on Communication, Networks and Satellite (COMNETSAT) (pp. 80–84).

Heyer, L. J., Kruglyak, S., & Yooseph, S. (1999). Expression Data : Identification and Analysis of Coexpressed Genes. Genome Research, (213), 1106–1115. https://doi.org/10.1101/gr.9.11.1106

Horng, S., Su, M., Chen, Y.-H., Kao, T., Chen, R., Lai, J., & Perkasa, C. D. (2011). A novel intrusion detection system based on hierarchical clustering and support vector machines.

Expert Systems with Applications, 38(1), 306–313. https://doi.org/10.1016/j.eswa.2010.06.066

Kim, S. (2012). Compute Spearman Correlation Coefficient with Matlab / CUDA. In Signal Processing and Information Technology (ISSPIT), 2012 IEEE International Symposium on (pp. 55–60).

Muchammad, K., & Ahmad, T. (2015). Detecting Intrusion Using Recursive Clustering and Sum of Log Distance to Sub-centroid. In Procedia - Procedia Computer Science (Vol. 72, pp. 446–452). Elsevier Masson SAS. https://doi.org/10.1016/j.procs.2015.12.125

Muttaqien, I. Z., & Ahmad, T. (2016). Increasing Performance of IDS by Selecting and Transforming Features. In 2016 IEEE International Conference on Communication, Networks and Satellite (COMNETSAT) Using (pp. 85–90).

Nasiroh Omar. (2014) Modelling Complexities Of Learner’s In Handling Web Texts Via Abstract Scene Analysis, Malaysian Journal of Computing, 2(1), 13-26.

Sembiring, R. W., Zain, J.M., and Embong, A. (2010). A Comparative Agglomerative Hierarchical Clustering Method to Cluster Implemented Course, Journal of Computing, Vol. 2, Issue 12 , 33-38.

Shiravi, A., Shiravi, H., Tavallaee, M., & Ghorbani, A. A. (2012). Toward developing a systematic approach to generate benchmark datasets for intrusion detection. Computers and Security, 31(3), 357–374. https://doi.org/10.1016/j.cose.2011.12.012

Published

2018-12-01

How to Cite

Gervais Hatungimana. (2018). PAIRWISE CLUSTERS OPTIMIZATION AND CLUSTER MOST SIGNIFICANT FEATURE METHODS FOR ANOMALY-BASED NETWORK INTRUSION DETECTION SYSTEM (POC2MSF). Malaysian Journal of Computing, 3(2). https://doi.org/10.24191/mjoc.v3i2.3598