Machine Learning-Based Data Sampling Techniques for Big Data Analytics
-
DOI:
https://doi.org/10.67228/3142788X/IJMLPA-2023PII3Q7VPublished 07-03-2023
Big Data Analytics, Machine Learning, Intelligent Sampling, Data Mining, Distributed Computing, Adaptive Sampling, Feature Selection, Hadoop, Apache Spark, Predictive Analytics, Reinforcement Learning, Deep Learning, Statistical Sampling, Large-Scale Data Processing Issue
Section
ArticlesHow to Cite
[1]C. Mendes, “Machine Learning-Based Data Sampling Techniques for Big Data Analytics”, IJMLPA, vol. 6, no. 2, pp. 01–17, Jul. 2023, doi: 10.67228/3142788X/IJMLPA-2023PII3Q7V.Abstract
The rapid expansion of IoT, cloud computing, social media, and distributed systems has generated massive amounts of data, making big data analytics increasingly important. Traditional sampling methods are often insufficient for handling the complexity and scalability challenges of modern big data environments. To overcome these limitations, machine learning-based sampling techniques have been developed to intelligently select representative data while preserving analytical accuracy. This paper surveys supervised, unsupervised, reinforcement, active, and deep learning-based sampling approaches and their integration with platforms like Hadoop and Apache Spark. Experimental findings show that intelligent sampling improves scalability, accuracy, efficiency, and resource utilization. The study concludes that machine learning-based sampling is a key technology for future big data analytics, with promising directions in hybrid, federated, explainable, and real-time adaptive sampling models.
References
[1] The Art of Computer Programming, D. E. Knuth, The Art of Computer Programming: Seminumerical Algorithms, 3rd ed. Boston, MA, USA: Addison-Wesley, 1997.
[2] William G. Cochran, W. G. Cochran, Sampling Techniques, 3rd ed. New York, NY, USA: Wiley, 1977.
[3] Rajeev Motwani and Prabhakar Raghavan, “Randomized Algorithms for Massive Data Sets,” Communications of the ACM, vol. 38, no. 11, pp. 32–44, 1995.
[4] Jeffrey Scott Vitter, “Random Sampling with a Reservoir,” ACM Transactions on Mathematical Software, vol. 11, no. 1, pp. 37–57, 1985.
[5] Jiawei Han, Micheline Kamber, and Jian Pei, Data Mining: Concepts and Techniques, 3rd ed. Burlington, MA, USA: Morgan Kaufmann, 2011.
[6] Tom M. Mitchell, Machine Learning. New York, NY, USA: McGraw-Hill, 1997.
[7] Christopher M. Bishop, Pattern Recognition and Machine Learning. New York, NY, USA: Springer, 2006.
[8] Leo Breiman, “Random Forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001.
[9] Corinna Cortes and Vladimir Vapnik, “Support-Vector Networks,” Machine Learning, vol. 20, no. 3, pp. 273–297, 1995.
[10] J. MacQueen, “Some Methods for Classification and Analysis of Multivariate Observations,” in Proc. 5th Berkeley Symp. Mathematical Statistics and Probability, Berkeley, CA, USA, 1967, pp. 281–297.
[11] Martin Ester, Hans-Peter Kriegel, Jörg Sander, and Xiaowei Xu, “A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise,” in Proc. 2nd Int. Conf. Knowledge Discovery and Data Mining, Portland, OR, USA, 1996, pp. 226–231.
[12] Burr Settles, “Active Learning Literature Survey,” University of Wisconsin–Madison, Madison, WI, USA, Tech. Rep. 1648, 2010.
[13] Yann LeCun, Yoshua Bengio, and Geoffrey Hinton, “Deep Learning,” Nature, vol. 521, no. 7553, pp. 436–444, 2015.
[14] Sepp Hochreiter and Jürgen Schmidhuber, “Long Short-Term Memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997.
[15] Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, 2nd ed. Cambridge, MA, USA: MIT Press, 2018.
Downloads
How to Cite
[1]C. Mendes, “Machine Learning-Based Data Sampling Techniques for Big Data Analytics”, IJMLPA, vol. 6, no. 2, pp. 01–17, Jul. 2023, doi: 10.67228/3142788X/IJMLPA-2023PII3Q7V.