A Comparative Study of Supervised Learning Algorithms for High-Dimensional Data

  • Authors

    • Sanjay Verma Operations Manager, HCL Technologies, India. Author
    • Naveen Kumar Product Manager, Zoho Corporation, India. Author

    DOI:

    https://doi.org/10.67228/3142788X/IJMLPA-2023PI3R7H

    Published 01-02-2023

  • High-Dimensional Data, Supervised Learning, Classification Algorithms, Feature Selection, Curse Of Dimensionality, Machine Learning Evaluation

    Issue

    Section

    Articles

    How to Cite

    [1]
    S. Verma and N. Kumar, “A Comparative Study of Supervised Learning Algorithms for High-Dimensional Data”, IJMLPA, vol. 6, no. 1, pp. 01–14, Jan. 2023, doi: 10.67228/3142788X/IJMLPA-2023PI3R7H.
  • Abstract

    High-dimensional data are now ubiquitous in the modern science and industry, such as bioinformatics, text mining, computer vision, finance, and cybersecurity. A prominent feature of such data is having many features in comparison with the number of observations, which may cause the judgement problem of the curse of dimensionality, greater computational cost, feature overlap, and overfitting. Though supervised learning algorithms are extensively used to do predictive modeling, they have very different performance properties in high dimensional feature space. The paper contains a thorough comparison of some of the most popular supervised learning algorithms in the high-dimensional data analysis scenario. The paper provides a systematic comparison between the linear, non-linear, probabilistic, and ensemble-based classifiers, which are: Logistic Regression, Support Vector Machine, k -Nearest Neighbor, Decision Tree, Random Forest, Naive Bayes, and Artificial Neural Network. Special attention is given to the study of the behaviour of an algorithm based on scalability, ability to generalize, resistance to noise, feature sparsity, and interpretability. Besides, the paper explores how dimensionality reduction and feature selection methods impact on the performance of classification. It suggests a single experimental procedure with standardized preprocessing pipelines, cross-validation schemes and performance metrics accuracy, precision, recall, F1-score and cost of the computation. To give the concept theoretical background, mathematical formulations of learning objectives and decision functions are given. The comparative analysis indicates that there is no universal algorithm that has the best performance in all high-dimensional conditions; the performance highly depends on the sample size, the features correlation, the level of data distribution as well as noise. This study has practical implications on researchers and practitioners to consider the proper supervised learning model to use the high-dimensional datasets and identifies future research opportunities in scalable and interpretable learning.

  • References

    [1] R. Bellman, Adaptive Control Processes: A Guided Tour, Princeton, NJ, USA: Princeton Univ. Press, 1961.

    [2] D. L. Donoho, “High-dimensional data analysis: The curses and blessings of dimensionality,” AMS Math Challenges Lecture, vol. 1, pp. 1–32, 2000.

    [3] T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd ed., New York, NY, USA: Springer, 2009.

    [4] J. Friedman, “Regularized discriminant analysis,” J. Amer. Statist. Assoc., vol. 84, no. 405, pp. 165–175, 1989.

    [5] R. Tibshirani, “Regression shrinkage and selection via the Lasso,” J. Roy. Statist. Soc. B, vol. 58, no. 1, pp. 267–288, 1996.

    [6] A. E. Hoerl and R. W. Kennard, “Ridge regression: Biased estimation for nonorthogonal problems,” Technometrics, vol. 12, no. 1, pp. 55–67, 1970.

    [7] V. Vapnik, The Nature of Statistical Learning Theory, 2nd ed., New York, NY, USA: Springer, 1999.

    [8] C. Cortes and V. Vapnik, “Support-vector networks,” Mach. Learn., vol. 20, no. 3, pp. 273–297, 1995.

    [9] T. Joachims, “Text categorization with support vector machines: Learning with many relevant features,” in Proc. ECML, Berlin, Germany: Springer, 1998, pp. 137–142.

    [10] A. McCallum and K. Nigam, “A comparison of event models for Naive Bayes text classification,” in Proc. AAAI Workshop on Learning for Text Categorization, 1998, pp. 41–48.

    [11] D. J. C. MacKay, Information Theory, Inference, and Learning Algorithms, Cambridge, U.K.: Cambridge Univ. Press, 2003.

    [12] L. Breiman, “Random forests,” Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001.

    [13] G. Biau and E. Scornet, “A random forest guided tour,” TEST, vol. 25, no. 2, pp. 197–227, 2016.

    [14] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, pp. 436–444, 2015.

    [15] I. Guyon and A. Elisseeff, “An introduction to variable and feature selection,” J. Mach. Learn. Res., vol. 3, pp. 1157–1182, 2003.

  • Downloads