Optimized Data Lake Architectures for Scalable ML Workflows

  • Authors

    • Sergey Lebedev Director, Institute of Precision Mechanics, Ukraine. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-2023PII2W4D

    Published 09-03-2023

  • Data Lake Architecture, Scalable Machine Learning, Data Engineering, Delta Lake, Apache Iceberg, Feature Store, Metadata Management, ML Pipeline, Optimization, Distributed Compute, Lakehouse

    Issue

    Section

    Articles

    How to Cite

    [1]
    S. Lebedev, “Optimized Data Lake Architectures for Scalable ML Workflows”, IJDEIC, vol. 6, no. 2, pp. 01–13, Sep. 2023, doi: 10.67228/30715717/IJDEIC-2023PII2W4D.
  • Abstract

    The exponential growth of data and the increasing demand for real-time, large-scale machine learning (ML) applications have challenged traditional data storage and processing architectures. Data lakes have emerged as a flexible and cost-effective solution for storing massive amounts of heterogeneous data. However, without optimization, data lakes can become inefficient, leading to performance bottlenecks in ML workflows. This paper presents an in-depth exploration of optimized data lake architectures tailored for scalable ML pipelines. We examine architectural best practices, including the separation of storage and compute, data versioning, metadata management, and the use of open table formats like Delta Lake, Apache Iceberg, and Apache Hudi. Furthermore, we propose a reference architecture that integrates modern orchestration tools, feature stores, and distributed compute engines to streamline the ML lifecycle from data ingestion to model deployment. Performance benchmarks and use cases demonstrate the proposed architecture's effectiveness in improving scalability, maintainability, and end-to-end ML throughput.

  • References

    [1] Armbrust; M.; et al. "Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores." Proceedings of VLDB Endowment; 2020.

    [2] Venkataraman; S.; et al. "Ernest: Efficient performance prediction for large-scale advanced analytics." USENIX Symposium on Networked Systems Design and Implementation (NSDI); 2016.

    [3] Netflix Technology Blog. “Introducing Iceberg: Netflix’s High-Performance Table Format for Huge Analytic Datasets.” 2019. https://netflixtechblog.com/

    [4] Uber Engineering. “Hudi: Uber’s Incremental Processing Framework for Big Data.” 2018. https://eng.uber.com/hudi/

    [5] Zaharia; M.; et al. "Apache Spark: A Unified Engine for Big Data Processing." Communications of the ACM; 2016.

    [6] Feast: Feature Store for Machine Learning. Open Source Documentation. https://docs.feast.dev/

    [7] Databricks. “The Data Lakehouse: A New Generation of Open Platforms That Unify Data Warehousing and Advanced Analytics.” White Paper; 2021.

    [8] Amazon Web Services. “AWS Lake Formation: Build a secure data lake in days.” AWS Whitepaper; 2022.

    [9] Kleppmann; M. Designing Data-Intensive Applications. O'Reilly Media; 2017.

    [10] Google Cloud. “BigLake: Unifying Data Warehousing and Data Lakes.” Google Cloud Blog; 2022.

    [11] Hopsworks. “The Case for a Feature Store in Modern ML Pipelines.” White Paper; 2021. https://www.hopsworks.ai/

  • Downloads