Scalable Data Engineering Architectures for Federated Learning in Decentralized Cloud Environments

  • Authors

    • Laura Conti Business Analyst, Deloitte Italy, Italy. Author
    • Dr. Andrew Collins Professor, University of Oxford, UK. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-2021PII4V3D

    Published 10-05-2021

  • Federated Learning, Decentralized Cloud, Scalable Data Engineering, Edge Computing, Data Pipeline Orchestration, Privacy-preserving Machine Learning, Kubernetes, Data Heterogeneity, Distributed Systems, Secure Data Sharing

    Issue

    Section

    Articles

    How to Cite

    [1]
    L. Conti and A. Collins, “Scalable Data Engineering Architectures for Federated Learning in Decentralized Cloud Environments”, IJDEIC, vol. 4, no. 2, pp. 01–12, Oct. 2021, doi: 10.67228/30715717/IJDEIC-2021PII4V3D.
  • Abstract

    The growing adoption of Federated Learning (FL) is reshaping the way machine learning models are trained across distributed, privacy-sensitive datasets. However, the scalable and efficient orchestration of data engineering pipelines in decentralized cloud environments remains a significant challenge. This paper presents a comprehensive architectural framework for scalable data engineering tailored for FL in heterogeneous and resource-constrained environments. By integrating modern distributed computing paradigms, such as Kubernetes-based orchestration, edge-aware data preprocessing, and secure federated communication, we propose a modular architecture that addresses data heterogeneity, scalability, and compliance. A case study in a healthcare IoT scenario validates the performance and flexibility of the proposed system. Our work serves as a blueprint for deploying robust FL systems in real-world decentralized cloud ecosystems.

  • References

    [1] McMahan, B., Moore, E., Ramage, D., Hampson, S., & Aguera y Arcas, B. (2017). Communication-efficient learning of deep networks from decentralized data. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS 2017) (pp. 1273–1282). PMLR.

    [2] Konečný, J., McMahan, H. B., Yu, F. X., Richtárik, P., Suresh, A. T., & Bacon, D. (2016). Federated learning: Strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492.

    [3] Smith, V., Chiang, C. K., Sanjabi, M., & Talwalkar, A. (2017). Federated multi-task learning. In Advances in Neural Information Processing Systems 30 (pp. 4424–4434).

    [4] Bonawitz, K., Ivanov, V., Kreuter, B., Marcedone, A., McMahan, H. B., Patel, S., Ramage, D., Segal, A., & Seth, K. (2017). Practical secure aggregation for privacy-preserving machine learning. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security (pp. 1175–1191).

    [5] Li, T., Sahu, A. K., Talwalkar, A., & Smith, V. (2018). Federated optimization in heterogeneous networks. arXiv preprint arXiv:1812.06127.

    [6] Hard, A., Rao, K., Mathews, R., Ramaswamy, S., Beaufays, F., Augenstein, S., Eichner, H., Kiddon, C., & Ramage, D. (2018). Federated learning for mobile keyboard prediction. arXiv preprint arXiv:1811.03604.

    [7] Bonawitz, K., Eichner, H., Grieskamp, W., Huba, D., Ingerman, A., Ivanov, V., Kiddon, C., Konečný, J., Mazzocchi, S., McMahan, H. B., Van Overveldt, T., Petrou, D., Ramage, D., & Roselander, J. (2019). Towards federated learning at scale: System design. In Proceedings of the 2nd MLSys Conference.

    [8] Satyanarayanan, M. (2017). The emergence of edge computing. Computer, 50(1), 30–39. https://doi.org/10.1109/MC.2017.9

    [9] Shi, W., Cao, J., Zhang, Q., Li, Y., & Xu, L. (2016). Edge computing: Vision and challenges. IEEE Internet of Things Journal, 3(5), 637–646. https://doi.org/10.1109/JIOT.2016.2579198

    [10] Mao, Y., You, C., Zhang, J., Huang, K., & Letaief, K. B. (2017). A survey on mobile edge computing: The communication perspective. IEEE Communications Surveys & Tutorials, 19(4), 2322–2358. https://doi.org/10.1109/COMST.2017.2745201

    [11] Varghese, B., & Buyya, R. (2018). Next generation cloud computing: New trends and research directions. Future Generation Computer Systems, 79, 849–861. https://doi.org/10.1016/j.future.2017.09.020

    [12] Dean, J., & Ghemawat, S. (2008). MapReduce: Simplified data processing on large clusters. Communications of the ACM, 51(1), 107–113. https://doi.org/10.1145/1327452.1327492

    [13] Zaharia, M., Chowdhury, M., Franklin, M. J., Shenker, S., & Stoica, I. (2010). Spark: Cluster computing with working sets. In Proceedings of the 2nd USENIX Conference on Hot Topics in Cloud Computing (pp. 10–10).

    [14] Gubbi, J., Buyya, R., Marusic, S., & Palaniswami, M. (2013). Internet of Things (IoT): A vision, architectural elements, and future directions. Future Generation Computer Systems, 29(7), 1645–1660. https://doi.org/10.1016/j.future.2013.01.010

    [15] Lim, W. Y. B., Luong, N. C., Hoang, D. T., Jiao, Y., Liang, Y. C., Yang, Q., Niyato, D., & Miao, C. (2019). Federated learning in mobile edge networks: A comprehensive survey. arXiv preprint arXiv:1909.11875.

  • Downloads