Self-Supervised Learning Techniques for Large-Scale AI Systems
-
DOI:
https://doi.org/10.67228/3071561X/IJIRHT-2023PII4P6KPublished 09-05-2023
Self-Supervised Learning, Contrastive Learning, Large-Scale AI, Representation Learning, Deep Learning, Unlabeled Data, Transformer Models, Pretraining Issue
Section
ArticlesHow to Cite
[1]S. Gupta, “Self-Supervised Learning Techniques for Large-Scale AI Systems”, IJIRHT, vol. 6, no. 2, pp. 01–11, Sep. 2023, doi: 10.67228/3071561X/IJIRHT-2023PII4P6K.Abstract
Self-supervised learning (SSL) is transforming artificial intelligence by enabling models to learn from large amounts of unlabeled data. Instead of relying on manual annotations, SSL leverages inherent data patterns to generate pseudo-labels, making it highly scalable and efficient for modern AI systems. Techniques such as contrastive learning, masked modeling, generative pretraining, and clustering have shown strong performance across vision, language, and speech tasks. This study examines key SSL methods, architectures, and training strategies, while addressing challenges like computational cost, feature collapse, and data bias. It proposes a unified framework that combines contrastive and generative approaches for improved efficiency and representation learning. Experimental results demonstrate that SSL outperforms traditional supervised learning in accuracy, scalability, and transferability, while also reducing data labeling costs. Future directions include integrating multimodal, reinforcement, and continual learning to further enhance SSL systems.
References
[1] Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). A Simple Framework for Contrastive Learning of Visual Representations (SimCLR). ICML.
[2] He, K., Fan, H., Wu, Y., Xie, S., & Girshick, R. (2020). Momentum Contrast for Unsupervised Visual Representation Learning (MoCo). CVPR.
[3] Grill, J. B., et al. (2020). Bootstrap Your Own Latent (BYOL): A New Approach to Self-Supervised Learning. NeurIPS.
[4] Caron, M., et al. (2020). Unsupervised Learning of Visual Features by Contrasting Cluster Assignments (SwAV). NeurIPS.
[5] Caron, M., Bojanowski, P., Joulin, A., & Douze, M. (2018). Deep Clustering for Unsupervised Learning of Visual Features (DeepCluster). ECCV.
[6] Kingma, D. P., & Welling, M. (2014). Auto-Encoding Variational Bayes (VAE). ICLR.
[7] Vincent, P., et al. (2008). Extracting and Composing Robust Features with Denoising Autoencoders. ICML.
[8] Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL.
[9] He, K., Chen, X., Xie, S., Li, Y., Dollár, P., & Girshick, R. (2022). Masked Autoencoders Are Scalable Vision Learners (MAE). CVPR.
[10] Dosovitskiy, A., et al. (2021). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT). ICLR.
[11] Zbontar, J., et al. (2021). Barlow Twins: Self-Supervised Learning via Redundancy Reduction. ICML.
[12] Chen, X., & He, K. (2021). Exploring Simple Siamese Representation Learning (SimSiam). CVPR.
[13] Misra, I., & van der Maaten, L. (2020). Self-Supervised Learning of Pretext-Invariant Representations (PIRL). CVPR.
[14] Jing, L., & Tian, Y. (2020). Self-Supervised Visual Feature Learning with Deep Neural Networks: A Survey. IEEE TPAMI.
[15] Ericsson, L., et al. (2021). How Well Do Self-Supervised Models Transfer? CVPR.
Downloads
How to Cite
[1]S. Gupta, “Self-Supervised Learning Techniques for Large-Scale AI Systems”, IJIRHT, vol. 6, no. 2, pp. 01–11, Sep. 2023, doi: 10.67228/3071561X/IJIRHT-2023PII4P6K.