HANDLING HIGHLY IMBALANCED OUTPUT CLASS LABEL: A CASE STUDY ON FANTASY PREMIER LEAGUE (FPL) VIRTUAL PLAYER PRICE CHANGES PREDICTION USING MACHINE LEARNING

Authors

  • Muhammad Muhaimin Khamsan Faculty of Computer and Mathematical Sciences, Universiti Teknologi MARA (UiTM) Shah Alam, Selangor, Malaysia
  • Ruhaila Maskat Faculty of Computer and Mathematical Sciences, Universiti Teknologi MARA (UiTM) Shah Alam, Selangor, Malaysia

DOI:

https://doi.org/10.24191/mjoc.v4i2.7021

Keywords:

Imbalanced Class Label, SMOTE upsampling, Machine Learning, Price Changes Prediction

Abstract

In practice, a balanced target class is rare. However, an imbalanced target class can be handled by resampling the original dataset, either by oversampling/upsampling or undersampling/downsampling. A popular upsampling technique is Synthetic Minority Over-sampling Technique (SMOTE). This technique increases the minority class by generating synthetic class labels and assigned the class based on the K-Nearest Neighbour (K-NN). SMOTE upsampling can only upsample at most one minority class at a time, which means for a multiclass dataset, it needs to undergo multilayer SMOTE to balance the class label distribution. This paper aims to find a suitable method in handling imbalanced class using dataset from Fantasy Premier League (FPL) virtual player to predict price changes. The cleaned dataset has a highly imbalanced class distribution, where the frequency of “Price Remain Unchanged (PRU)” is higher than “Price Fall (PF)” and “Price Rise (PR)”. This paper compared between the baseline (original) dataset, SMOTE-applied dataset and shuffled, linear and stratified sampling in split train-test subset, based on a deep learning algorithm. This paper also proposed criteria of low values in standard deviation (distribution of true positive on each class label on accuracy) as a measurement for finding the best method in handling imbalanced class labels. As a result, multilayer SMOTE until all the classes distribution is the same, combined with stratified sampling in split training and testing subset, get the lower standard deviation (5.7873), high accuracy (80.06%) and less execution runtime (1 minute 41 seconds) compared to the original highly imbalanced dataset.

References

Alejo. R., Garcia. V., Mollineda . R. A., Sanchez. J. S., & Sotoca. J. M. (2007). The class imbalance problem in pattern classification and learning. Dept de Llenguatjes i Sistemes Informatics, Universitat Jaume I, Spain.

Anand. V. (2018). Fantasy Premier League dataset from season 2016/2017 [Data file]. Retrieved from Github: https://github.com/vaastav/Fantasy-Premier-League/tree/master/data/2016-17/gws

Barandela. R., Garcia. V., Rangel. E., & Sanchez. J. S. (2002). Strategies for learning in class imbalance problems. The journal of the recognition society, 849-851.

Benabbou. F., Sadgali. A., & Sael. N. (2019). Performance of machine learning techniques in the detection of financial frauds. Second International Conference on Intelligent Computing in Data Sciences (ICDS 2018) (pp. 45-54). Morocco: Elsevier.

Bowyer. K. W., Chawla. N. V., Hall. L.O., & Kegelmeyer, W. P. (2002). SMOTE : Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, vol 16, 321-357. AI Access Foundation and Morgan Kaufmann Publishers.

Chung. J. Y & Lee. S. (2019). Droupout early warning system for high school students using machine learning. Children and Youth Services Review, 346-353.

Glauner. P., State. R., & Valtchev. P. (2017). Impact of Biases in Big Data. Luxembourg: National Research Fund.

He. C., Jiang. X., Xiao. J., & Xie. L. (2011). Dynamic classifier ensemble model for customer classification with imbalance class distribution. Expert System with Application 39 (2012), Elsevier Ltd.

Helmenstine. A. M. (2018, September 27). How to calculate population standard deviation. Retrieved from ThoughtCo.: https://www.thoughtco.com/population-standard-deviation-calculation-609522

Ho. S., Ng. M. K., Wu. Q., Ye. Y., & Zhang. H. (2014). ForesTexter: An efficient random forest algorithm for imbalanced text categorization. Knowledge Based System 67 (2014), Elsevier B. V.

Kaur. H., & Kumari. V. (2018). Predictive modelling and analytics for diabetes using a machine learning approach. Applied Computing and Informatics.

Kubat. M., & Matwin. S. (1997). Addressing the Curse of Imbalanced Training Sets: One sided Selection. Proceedings of the 14th International Conference on Machine Learning (pp. 179-186). Nashviille USA: University of Ottawa.

Liu. J., Luo. X., Tang. Y., Xu. Z., Yang. Z., Yuan. P., Zhang. T., & Zhang. Y. (2019). Software defect prediction based on kernel PCA and weighted extreme learning machine. Inforrmation and Software Technology, 182-200.

Ma. Z., Wang. G., Wang. Z., Xue. J., & Zhu. R. (2018). LRID : A new metric of multi-class imblance degree based on likelihood-ratio test. Pattern Recognition Letters 116 (2018), 36-42.

Rocca. B. (2018, January 28). Handling imblanced datasets in machine learning. Retrieved from Towads Data Science: https://towardsdatascience.com/handling-imbalanced-datasets-in-machine-learning-7a0e84220f28

Severin. E., & Veganzones. D. (2018). An investigation of bankruptcy prediction in imbalanced datasets. Decision Support Systems 112 (2018), 111-124.

Stapor. K. (2018). Evaluating and Comparing Classifiers: Review, Some Recommendations and Limitations. Proceedings of the 10th International Conference on Computer Recognition System CORES 2017. Advanced in Intelligent and Computing, vol 578. Springer.

Published

2019-12-01

How to Cite

Muhammad Muhaimin Khamsan, & Ruhaila Maskat. (2019). HANDLING HIGHLY IMBALANCED OUTPUT CLASS LABEL: A CASE STUDY ON FANTASY PREMIER LEAGUE (FPL) VIRTUAL PLAYER PRICE CHANGES PREDICTION USING MACHINE LEARNING. Malaysian Journal of Computing, 4(2). https://doi.org/10.24191/mjoc.v4i2.7021