Unified Data Lake Architectures: Integrating Object-to-Stream Ingestion, Server less Kafka Streaming ETL, and Near Real-Time Analytical Indexing

  • Authors

    • Mahesh Kumar Goyal Senior Data Architect, Amazon LLC, USA. Author
    • Suresh Patnam Senior Data Architect, Amazon LLC, USA. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-2020PI7P2M

    Published 06-13-2020

  • Data Lake, Change Data Capture (CDC), Amazon Kinesis, AWS Glue ETL, Apache Parquet, Elasticsearch, Amazon S3, Serverless Analytics

    Issue

    Section

    Articles

    How to Cite

    [1]
    M. K. Goyal and S. Patnam, “Unified Data Lake Architectures: Integrating Object-to-Stream Ingestion, Server less Kafka Streaming ETL, and Near Real-Time Analytical Indexing”, IJDEIC, vol. 3, no. 1, pp. 01–08, Jun. 2020, doi: 10.67228/30715717/IJDEIC-2020PI7P2M.
  • Abstract

    Enterprise data platforms often run into a common operational dilemma: business stakeholders require sub-second operational log search, while data engineering and BI teams require massive, multi-terabyte analytical queries over historical data. Building separate streaming and batch infrastructure leads to duplicate codebases, out-of-order event drift, and excessive cloud infrastructure costs. In this paper, we describe a unified, production-tested cloud data lake architecture that solves these challenges on AWS. Instead of running periodic, heavy SQL queries against production transactional databases, we repurpose AWS Database Migration Service (DMS) as a permanent, non-blocking Change Data Capture (CDC) streaming source that captures row-level binary log deltas and pushes them directly into Amazon Kinesis Data Streams and Amazon S3. For downstream analytical processing, we implement an adaptive serverless Apache Spark ETL pipeline using AWS Glue to compact micro-batches into 256 MB Snappy-compressed Apache Parsquet partitions, resolving the common 'small file problem' on object storage. In parallel, operational server access logs are routed via event-driven AWS Lambda consumers to Amazon Elasticsearch Service (7.x) using a tri-tier index lifecycle management strategy. We evaluated this architecture across real production enterprise workloads sustaining over 100,000 transactions per second (TPS). Our benchmarks show a 78.4% reduction in median data ingestion latency (p99 < 4.2s), an 82.8% reduction in S3 storage footprints via columnar compression, and a 19x speedup in distributed SQL query execution compared to raw object scans.

  • References

    [1] N. Marz and J. Warren, Big Data: Principles and best practices of scalable realtime data systems. Manning Publications, 2015.

    [2] J. Kreps, N. Narkhede, and J. Rao, "Kafka: A distributed messaging system for log processing," in Proceedings of the NetDB Workshop, 2011, pp. 1–7.

    [3] M. Zaharia et al., "Apache Spark: A unified engine for big data processing," Communications of the ACM, vol. 59, no. 11, pp. 56–65, 2016.

    [4] T. Akidau et al., "The Dataflow Model: A practical approach to balancing correctness, latency, and cost in massive-scale, unbounded, out-of-order data processing," Proceedings of the VLDB Endowment, vol. 8, no. 12, pp. 1792–1803, 2015.

    [5] S. Melnik et al., "Dremel: Interactive analysis of web-scale datasets," Proceedings of the VLDB Endowment, vol. 3, no. 1-2, pp. 330–339, 2010.

    [6] C. Gormley and Z. Tong, Elasticsearch: The Definitive Guide. O'Reilly Media, 2015.

    [7] K. Shvachko, H. Kuang, S. Radia, and R. Chansler, "The Hadoop Distributed File System," in 2010 IEEE 26th Symposium on Mass Storage Systems and Technologies (MSST), 2010, pp. 1–10.

    [8] P. Carbone et al., "Apache Flink: Stream and batch processing in a single engine," IEEE Data Engineering Bulletin, vol. 38, no. 4, pp. 28–38, 2015.

    [9] J. Dean and S. Ghemawat, "MapReduce: Simplified data processing on large clusters," Communications of the ACM, vol. 51, no. 1, pp. 107–113, 2008.

    [10] Amazon Web Services, AWS Database Migration Service User Guide. Amazon.com, Inc., 2020.

    [11] Amazon Web Services, AWS Glue Developer Guide: Developing ETL Scripts and Streaming Jobs. Amazon.com, Inc., 2020.

  • Downloads