AI-Driven Data Sampling Strategies for Efficient Model Training
-
DOI:
https://doi.org/10.67228/3142788X/IJMLPA-2024PI7T2CPublished 01-03-2024
Artificial Intelligence, Data Sampling, Machine Learning, Active Learning, Importance Sampling, Model Training, Big Data Analytics, Deep Learning, Computational Efficiency, Intelligent Data Selection Issue
Section
ArticlesHow to Cite
[1]K. Balasubramanian and M. Krishnan, “AI-Driven Data Sampling Strategies for Efficient Model Training”, IJMLPA, vol. 7, no. 1, pp. 01–17, Jan. 2024, doi: 10.67228/3142788X/IJMLPA-2024PI7T2C.Abstract
Artificial Intelligence (AI) and Machine Learning (ML) are widely used across industries such as healthcare, finance, manufacturing, transportation, cybersecurity, and scientific computing to solve complex problems. The effectiveness of ML models depends heavily on the quality and representativeness of training data. However, the rapid growth of big data has increased computational costs, memory requirements, and training times, making traditional full-dataset training inefficient.To address these challenges, intelligent data sampling has emerged as an important research area. AI-based sampling techniques—including random sampling, stratified sampling, importance sampling, active learning, uncertainty sampling, reinforcement learning-based sampling, and adaptive data selection—identify the most informative data points while reducing redundancy. These methods improve computational efficiency, enhance class representation, and accelerate model training without significantly affecting accuracy. This study reviews the evolution of AI-driven sampling strategies and proposes an adaptive framework that combines statistical sampling with intelligent decision-making mechanisms for real-time data selection. Experimental results demonstrate that intelligent sampling can reduce training data requirements by 30–60% while maintaining over 95% of the predictive performance achieved using full datasets. The approach also improves scalability, reduces energy consumption, and supports efficient deployment of modern machine learning systems. Overall, AI-based sampling provides a practical and effective solution for optimizing data usage and enhancing machine learning performance in large-scale applications.
References
[1] D. D. Lewis and W. A. Gale, “A sequential algorithm for training text classifiers,” in Proceedings of the 17th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Dublin, Ireland, 1994, pp. 3–12.
[2] D. Cohn, L. Atlas, and R. Ladner, “Improving generalization with active learning,” Machine Learning, vol. 15, no. 2, pp. 201–221, 1994.
[3] H. S. Seung, M. Opper, and H. Sompolinsky, “Query by committee,” in Proceedings of the Fifth Annual Workshop on Computational Learning Theory, Pittsburgh, PA, USA, 1992, pp. 287–294.
[4] B. Settles, “Active learning literature survey,” University of Wisconsin-Madison, Computer Sciences Technical Report 1648, 2010.
[5] A. Beygelzimer, S. Dasgupta, and J. Langford, “Importance weighted active learning,” in Proceedings of the 26th Annual International Conference on Machine Learning, Montreal, Canada, 2009, pp. 49–56.
[6] Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in Proceedings of the 26th International Conference on Machine Learning, Montreal, Canada, 2009, pp. 41–48.
[7] O. Chapelle and L. Li, “An empirical evaluation of Thompson sampling,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 24, 2011, pp. 2249–2257.
[8] R. Sutton and A. Barto, Reinforcement Learning: An Introduction, 2nd ed. Cambridge, MA, USA: MIT Press, 2018.
[9] V. Mnih et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, pp. 529–533, 2015.
[10] T. Katharopoulos and F. Fleuret, “Not all samples are created equal: Deep learning with importance sampling,” in Proceedings of the 34th International Conference on Machine Learning (ICML), Sydney, Australia, 2017, pp. 2525–2534.
[11] A. Katharopoulos and F. Fleuret, “Biased importance sampling for deep neural network training,” arXiv preprint arXiv:1711.00043, 2017.
[12] S. Coleman, A. Bialkowski, and T. F. Gonzalez, “Sample-efficient adaptive training for deep neural networks,” IEEE Transactions on Neural Networks and Learning Systems, vol. 31, no. 9, pp. 3402–3415, 2020.
[13] C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in Proceedings of the 34th International Conference on Machine Learning (ICML), Sydney, Australia, 2017, pp. 1126–1135.
[14] H. H. Bui, S. Venkatesh, and G. West, “Policy-gradient reinforcement learning methods for adaptive data selection,” IEEE Transactions on Knowledge and Data Engineering, vol. 29, no. 5, pp. 1124–1137, 2017.
[15] T. Elsken, J. H. Metzen, and F. Hutter, “Neural architecture search: A survey,” Journal of Machine Learning Research, vol. 20, no. 55, pp. 1–21, 2019.
[16] B. Zoph and Q. V. Le, “Neural architecture search with reinforcement learning,” in International Conference on Learning Representations (ICLR), 2017.
[17] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of CVPR, 2016, pp. 770–778.
[18] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” Communications of the ACM, vol. 60, no. 6, pp. 84–90, 2017.
[19] A. Graves, M. G. Bellemare, J. Menick, R. Munos, and K. Kavukcuoglu, “Automated curriculum learning for neural networks,” in Proceedings of ICML, 2017.
[20] O. Sener and S. Savarese, “Active learning for convolutional neural networks: A core-set approach,” in International Conference on Learning Representations (ICLR), 2018.
[21] Gajula, S. (2023). A review of anomaly identification in finance frauds using machine learning system. International Journal of Current Engineering and Technology, 13(6), 568–575. https://ijcet.evegenis.org/index.php/ijcet/article/view/820
Downloads
How to Cite
[1]K. Balasubramanian and M. Krishnan, “AI-Driven Data Sampling Strategies for Efficient Model Training”, IJMLPA, vol. 7, no. 1, pp. 01–17, Jan. 2024, doi: 10.67228/3142788X/IJMLPA-2024PI7T2C.