Machine Learning Frameworks for Large-Scale Data Forecasting

  • Authors

    • Dr. John Peterson Professor, Stanford University, USA. Author

    DOI:

    https://doi.org/10.67228/3142788X/IJMLPA-2024PII2V9R

    Published 10-03-2024

  • Large-Scale Forecasting, Machine Learning Frameworks, Predictive Analytics, Big Data, Distributed Computing, Apache Spark, TensorFlow, Ensemble Learning, Forecasting Models, Data Mining

    Issue

    Section

    Articles

    How to Cite

    [1]
    J. Peterson, “Machine Learning Frameworks for Large-Scale Data Forecasting”, IJMLPA, vol. 7, no. 2, pp. 01–16, Oct. 2024, doi: 10.67228/3142788X/IJMLPA-2024PII2V9R.
  • Abstract

    Machine learning has become a critical technology for large-scale data forecasting across industries such as finance, healthcare, transportation, manufacturing, energy, and e-commerce. While traditional forecasting methods like linear regression, ARIMA, and exponential smoothing perform well on smaller datasets, they struggle to manage the volume, velocity, and complexity of modern big data. Machine learning frameworks such as TensorFlow, Apache Spark MLlib, H2O.ai, Scikit-learn, and XGBoost offer scalable solutions by processing large datasets, identifying complex patterns, and generating accurate predictions through distributed computing.This study reviews machine learning architectures for large-scale forecasting, covering forecasting evolution, key algorithms, framework architectures, data processing pipelines, and evaluation methods. It proposes a scalable forecasting framework consisting of data collection, preprocessing, feature engineering, model training, forecasting, and performance evaluation. Findings indicate that distributed machine learning frameworks improve forecasting accuracy while reducing computational costs. Ensemble learning and gradient boosting techniques demonstrate superior performance, scalability, and robustness compared to traditional methods. The study concludes that machine learning frameworks provide a strong foundation for large-scale forecasting and will play an increasingly important role in future predictive applications.

  • References

    [1] G. E. P. Box, G. M. Jenkins, G. C. Reinsel, and G. M. Ljung, Time Series Analysis: Forecasting and Control, 5th ed. Hoboken, NJ, USA: Wiley, 2015.

    [2] R. J. Hyndman and G. Athanasopoulos, Forecasting: Principles and Practice, 3rd ed. Melbourne, Australia: OTexts, 2021.

    [3] C. Chatfield, The Analysis of Time Series: An Introduction, 6th ed. Boca Raton, FL, USA: CRC Press, 2016.

    [4] J. H. Friedman, “Greedy function approximation: A gradient boosting machine,” Annals of Statistics, vol. 29, no. 5, pp. 1189–1232, 2001.

    [5] L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001.

    [6] V. N. Vapnik, The Nature of Statistical Learning Theory. New York, NY, USA: Springer, 1995.

    [7] T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd ed. New York, NY, USA: Springer, 2009.

    [8] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA, USA: MIT Press, 2016.

    [9] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, no. 7553, pp. 436–444, 2015.

    [10] T. White, Hadoop: The Definitive Guide, 4th ed. Sebastopol, CA, USA: O’Reilly Media, 2015.

    [11] M. Zaharia et al., “Apache Spark: A unified engine for big data processing,” Communications of the ACM, vol. 59, no. 11, pp. 56–65, 2016.

    [12] M. Abadi et al., “TensorFlow: A system for large-scale machine learning,” in Proc. 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI), Savannah, GA, USA, 2016, pp. 265–283.

    [13] S. Landset, T. M. Khoshgoftaar, A. N. Richter, and T. Hasanin, “A survey of open source tools for machine learning with big data in the Hadoop ecosystem,” Journal of Big Data, vol. 2, no. 1, pp. 1–36, 2015.

    [14] X. Wu, X. Zhu, G.-Q. Wu, and W. Ding, “Data mining with big data,” IEEE Transactions on Knowledge and Data Engineering, vol. 26, no. 1, pp. 97–107, 2014.

    [15] J. Manyika et al., Big Data: The Next Frontier for Innovation, Competition, and Productivity. San Francisco, CA, USA: McKinsey Global Institute, 2011.

    [16] Gajula, S. (2023). A review of anomaly identification in finance frauds using machine learning system. International Journal of Current Engineering and Technology, 13(6), 568–575. https://ijcet.evegenis.org/index.php/ijcet/article/view/820

  • Downloads