Data Lakes: Architectures and Systems for Big Data Management

  • Authors

    • Dr. Carlos Mendes Professor, University of Lisbon, Portugal. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-2019PI8L5T

    Published 06-12-2019

  • Data Lakes, Big Data Management, Metadata Management, Data Lake Architecture, Data Ingestion, Cloud Data Lakes, Apache Hadoop, Data Governance

    Issue

    Section

    Articles

    How to Cite

    [1]
    C. Mendes, “Data Lakes: Architectures and Systems for Big Data Management”, IJDEIC, vol. 2, no. 1, pp. 01–13, Jun. 2019, doi: 10.67228/30715717/IJDEIC-2019PI8L5T.
  • Abstract

    Data lakes have emerged as a fundamental paradigm for managing large-scale, heterogeneous, and high-velocity data generated in contemporary digital ecosystems. As enterprises increasingly transition from traditional data warehousing to more scalable and flexible data lake solutions, there is a growing need for standardized architectures and system models to ensure efficiency, reliability, and interoperability. This paper presents an in-depth analysis of data lake architectures, explores system components, and evaluates design considerations essential for handling big data. We survey existing frameworks and technologies, propose a reference architecture, and introduce a performance-oriented methodology for assessing system efficiency. Key challenges such as data ingestion, metadata management, data governance, security, and query performance are discussed with real-world examples and case studies. The paper further provides experimental results from implementing a prototype data lake using Apache Hadoop and Spark on cloud infrastructure. Performance metrics such as ingestion latency, query response time, and storage efficiency are analyzed. A comparative discussion of existing data lake solutions highlights their strengths and limitations, while the conclusion summarizes key insights and outlines future research directions for more intelligent and autonomous data lake systems.

  • References

    [1] Inmon, W. H. (2016). Data lake architecture: Designing the data lake and avoiding the garbage dump. Technics Publications.

    [2] Gorelik, A. (2016). The enterprise big data lake: Delivering the promise of big data and data science. O'Reilly Media.

    [3] Fang, H. (2015). Managing data lakes in big data era: What's a data lake and why has it become popular in data management ecosystem. In 2015 IEEE International Conference on Cyber Technology in Automation, Control, and Intelligent Systems (pp. 820–824). IEEE. https://doi.org/10.1109/CYBER.2015.7288049

    [4] Madera, C., & Laurent, A. (2016). The next information architecture evolution: The data lake wave. In Proceedings of the 8th International Conference on Management of Digital EcoSystems (pp. 174–180). ACM. https://doi.org/10.1145/3012071.3012077

    [5] Hai, R., Geisler, S., & Quix, C. (2016). Constance: An intelligent data lake system. In Proceedings of the 2016 International Conference on Management of Data (pp. 2097–2100). ACM. https://doi.org/10.1145/2882903.2899389

    [6] Thusoo, A., & Sharma, B. (2016). Architecting data lakes. O'Reilly Media.

    [7] Madsen, M. (2015). How to build an enterprise data lake: Important considerations before jumping. Third Nature Inc.

    [8] Zikopoulos, P., DeRoos, D., Bienko, C., Buglio, R., & Andrews, M. (2015). Big data beyond the hype: A guide to conversations for today's data center. McGraw-Hill Education.

    [9] Zaharia, M., Chowdhury, M., Das, T., Dave, A., Ma, J., McCauley, M., Franklin, M. J., Shenker, S., & Stoica, I. (2012). Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. In Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation (pp. 15–28).

    [10] White, T. (2015). Hadoop: The definitive guide (4th ed.). O'Reilly Media.

  • Downloads