Semi-Supervised Learning for Data-Scarce Environments
-
DOI:
https://doi.org/10.67228/3142788X/IJMLPA-2023PI4N7TPublished 06-04-2023
Semi-Supervised Learning, Data-Scarce Environments, Machine Learning, Self-Training, Pseudo-Labeling, Consistency Regularization, Deep Learning, Unlabeled Data Issue
Section
ArticlesHow to Cite
[1]A. Reza, “Semi-Supervised Learning for Data-Scarce Environments”, IJMLPA, vol. 6, no. 1, pp. 01–15, Jun. 2023, doi: 10.67228/3142788X/IJMLPA-2023PI4N7T.Abstract
In most practical machine learning systems, it is costly, time-consuming, and the process of obtaining labeled data is even impractical in certain areas like the medical field, remote sensing, cyber security and industrial automation, among others. Nonetheless, large volumes of raw data are usually easy to find. Semi-supervised learning (SSL) has become an impressive paradigm that applies both labeled and unlabeled data to enhance the learning performance in situations with scarce data. In this paper, we discuss a detailed research on semi-supervised learning methods in data-sparse setting and their theoretical basis, algorithm details and deployment strategies. We consider classical and deep learning-based methods of SSL, such as self-training, co-training, graph based, consistency regularization and pseudo-labeling. An integrated deployment approach is suggested regarding the implementation of the SSL systems in practice, where there are data constraints. The effectiveness of SSL methods in comparison to the completely supervised models is proved by a lot of experimental analysis and comparative evaluation. The findings show that semi-supervised learning provides significant improvement in accuracy of classification, robustness and generalization and also minimizes cost of annotation. The paper is concluded by the insights about the open research challenges and the further directions.
References
[1] Zhu, X. (2005).Semi-Supervised Learning Literature Survey. Computer Sciences Technical Report 1530, University of Wisconsin-Madison.
[2] Chapelle, O., Schölkopf, B., & Zien, A. (2006).Semi-Supervised Learning. MIT Press.
[3] Yarowsky, D. (1995). Unsupervised word sense disambiguation rivaling supervised methods. Proceedings of the 33rd Annual Meeting of the ACL.
[4] McCallum, A., & Nigam, K. (1998). Employing EM in text classification. AAAI Workshop on Text Classification.
[5] Blum, A., & Mitchell, T. (1998). Combining labeled and unlabeled data with co-training. Proceedings of the 11th Annual Conference on Computational Learning Theory (COLT).
[6] Zhou, D., Bousquet, O., Lal, T., Weston, J., & Schölkopf, B. (2004). Learning with local and global consistency. Advances in Neural Information Processing Systems (NeurIPS).
[7] Belkin, M., Niyogi, P., & Sindhwani, V. (2006). Manifold regularization: A geometric framework for learning from labeled and unlabeled examples. Journal of Machine Learning Research.
[8] Zhu, X., Ghahramani, Z., & Lafferty, J. (2003). Semi-supervised learning using Gaussian fields and harmonic functions. Proceedings of ICML.
[9] Lee, D.-H. (2013). Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. ICML Workshop on Challenges in Representation Learning.
[10] Laine, S., & Aila, T. (2017). Temporal ensembling for semi-supervised learning. International Conference on Learning Representations (ICLR).
[11] Sajjadi, M., Javanmardi, M., & Tasdizen, T. (2016). Regularization with stochastic transformations and perturbations for deep semi-supervised learning. Advances in Neural Information Processing Systems (NeurIPS).
[12] Tarvainen, A., & Valpola, H. (2017). Mean teachers are better role models. Advances in Neural Information Processing Systems (NeurIPS).
[13] Miyato, T., Maeda, S., Koyama, M., & Ishii, S. (2018). Virtual adversarial training: A regularization method for supervised and semi-supervised learning. IEEE Transactions on Pattern Analysis and Machine Intelligence.
[14] Sohn, K., et al. (2020). FixMatch: Simplifying semi-supervised learning with consistency and confidence. Advances in Neural Information Processing Systems (NeurIPS).
[15] Xie, Q., et al. (2020). Self-training with noisy student improves ImageNet classification. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
Downloads
How to Cite
[1]A. Reza, “Semi-Supervised Learning for Data-Scarce Environments”, IJMLPA, vol. 6, no. 1, pp. 01–15, Jun. 2023, doi: 10.67228/3142788X/IJMLPA-2023PI4N7T.