Enhancing Distributed Systems for Real-Time Machine Learning Model Deployment and Management
-
DOI:
https://doi.org/10.67228/30713315/IJAIDT-2020PI1A2KPublished 01-05-2020
Distributed Systems, Real-Time Processing, Machine Learning Deployment, Edge Computing, Federated Learning, System Scalability, Fault Tolerance, Dynamic Orchestration Issue
Section
ArticlesHow to Cite
[1]S. A. Rahman, “Enhancing Distributed Systems for Real-Time Machine Learning Model Deployment and Management”, IJAIDT, vol. 3, no. 1, pp. 01–07, Jan. 2020, doi: 10.67228/30713315/IJAIDT-2020PI1A2K.Abstract
The integration of machine learning (ML) models into distributed systems has become pivotal for applications requiring real-time data processing and decision-making. This paper investigates methodologies to enhance distributed architectures for the efficient deployment and management of ML models in real-time environments. We explore the challenges associated with latency, scalability, and fault tolerance, and propose solutions leveraging edge computing, federated learning, and dynamic orchestration. Through empirical evaluations, we demonstrate the efficacy of the proposed approaches in optimizing real-time ML workflows.
References
[1] Dean, J., & Ghemawat, S. (2008). MapReduce: Simplified Data Processing on Large Clusters. Communications of the ACM, 51(1), 107–113.
[2] Zaharia, M., et al. (2016). Apache Spark: A Unified Engine for Big Data Processing. Communications of the ACM, 59(11), 56–65.
[3] Li, M., et al. (2014). Scaling Distributed Machine Learning with the Parameter Server. OSDI.
[4] Abadi, M., et al. (2016). TensorFlow: A System for Large-Scale Machine Learning. OSDI.
[5] Moritz, P., et al. (2018). Ray: A Distributed Framework for Emerging AI Applications. OSDI.
[6] Crankshaw, D., et al. (2017). Clipper: A Low-Latency Online Prediction Serving System. NSDI.
[7] Olston, C., et al. (2017). TensorFlow-Serving: Flexible, High-Performance ML Serving. arXiv preprint arXiv:1712.06139.
[8] Kwon, Y., et al. (2019). SageMaker: Large-Scale Machine Learning in the Cloud. SIGMOD.
[9] Oakes, E., et al. (2020). SOCC’20 – Lighthouse: Flexible and Efficient ML Serving. ACM Symposium on Cloud Computing.
Downloads
How to Cite
[1]S. A. Rahman, “Enhancing Distributed Systems for Real-Time Machine Learning Model Deployment and Management”, IJAIDT, vol. 3, no. 1, pp. 01–07, Jan. 2020, doi: 10.67228/30713315/IJAIDT-2020PI1A2K.