Cloud-Native Data Pipelines for Enterprise Analytics
-
DOI:
https://doi.org/10.67228/30713498/IJADSMC-2020PI9F3RPublished 04-03-2020
Cloud-Native Architecture, Data Pipelines, Enterprise Analytics, Streaming Analytics, Microservices, Container Orchestration, Data Engineering, Scalability, Real-Time Processing Issue
Section
ArticlesHow to Cite
[1]E. Williams, “Cloud-Native Data Pipelines for Enterprise Analytics”, IJADSMC, vol. 3, no. 1, pp. 01–13, Apr. 2020, doi: 10.67228/30713498/IJADSMC-2020PI9F3R.Abstract
Cloud native data pipelines have become an enabling ingredient of the modern enterprise analytics to fulfill the ever-increasing demand of a scale-loving, resilient, and real-time processing of a wide range of data sources. Organizations currently produce large amounts of structured, semi-structured, and unstructured data in transactional systems, Internet of Things (IoT) platforms, digital channels, and data sources that are external (Bank of America 2017). Old monolithic data integration architectures are designed to provide batch-oriented processing and static infrastructure capabilities have challenges satisfying low latency, scale on demand, and 24/7 requirements. Reactively, cloud-native paradigms, including the foundations of microservices, container orchestration, event-driven architectures, and managed cloud services have caused a rethinking of the data pipeline design, deployment and operation. This article provides an in-depth analysis of cloud-native pipeline data to enterprise analytics along with their main architectural concepts and processing models as well as operational aspects that fall within the professional scope of IEEE publications. The research paper summarizes the literature and business methodologies to present a reference model which brings together data ingestion, stream processing, batch processing, storage, governance, and analytics consumption layers. Special concern is opened to the contributions of containerization, orchestration platforms, and serverless computing towards facilitation of elasticity and fault tolerance. The paper also examines design patterns like Lambda architecture and Kappa architecture, data mesh theory and metadata-based orchestration, with an emphasis on its application to the large enterprise environment. An organized approach to the design and deployment of cloud-native data pipelines with the inclusion of data quality management, security controls, observability, and cost optimization is suggested. Throughput, latency and scalability modeling mathematical formulations are proposed in order to facilitate capacity planning and performance measurement. Representative enterprise workloads as shown through experiment results exhibit evident increases in data processing latency, pipeline reliability and operational efficiency over traditional architectures. These findings are placed in context to the discussion of the broader transformation efforts at enterprises, whereas the conclusion provides recommendations on future research opportunities, such as autonomous pipeline optimization and AI-based orchestration.
References
[1] W. H. Inmon, Building the Data Warehouse, 4th ed. Hoboken, NJ, USA: Wiley, 2005.
[2] R. Kimball and M. Ross, The Data Warehouse Toolkit: The Definitive Guide to Dimensional Modeling, 3rd ed. Hoboken, NJ, USA: Wiley, 2013.
[3] J. Dean and S. Ghemawat, “MapReduce: Simplified data processing on large clusters,” Communications of the ACM, vol. 51, no. 1, pp. 107–113, Jan. 2008.
[4] M. Zaharia et al., “Spark: Cluster computing with working sets,” in Proc. 2nd USENIX Conf. Hot Topics in Cloud Computing (HotCloud), Boston, MA, USA, 2010, pp. 1–7.
[5] T. Akidau et al., “The Dataflow model: A practical approach to balancing correctness, latency, and cost in massive-scale, unbounded, out-of-order data processing,” Proc. VLDB Endowment, vol. 8, no. 12, pp. 1792–1803, Aug. 2015.
[6] J. Kreps, “Questioning the Lambda Architecture,” O’Reilly Radar, Jul. 2014.
[7] N. Marz and J. Warren, Big Data: Principles and Best Practices of Scalable Real-Time Data Systems. Shelter Island, NY, USA: Manning, 2015.
[8] S. Schneider, “Stream processing,” Communications of the ACM, vol. 59, no. 11, pp. 58–66, Nov. 2016.
[9] M. Fowler and J. Lewis, “Microservices: A definition of this new architectural term,” martinfowler.com, 2014.
[10] [10] B. Burns, B. Grant, D. Oppenheimer, E. Brewer, and J. Wilkes, “Borg, Omega, and Kubernetes,” Communications of the ACM, vol. 59, no. 5, pp. 50–57, May 2016.
[11] Balalaie, A. Heydarnoori, and P. Jamshidi, “Microservices architecture enables DevOps: Migration to a cloud-native architecture,” IEEE Software, vol. 33, no. 3, pp. 42–52, May–Jun. 2016.
[12] P. Jamshidi, C. Pahl, N. C. Mendonça, J. Lewis, and S. Tilkov, “Microservices: The journey so far and challenges ahead,” IEEE Software, vol. 35, no. 3, pp. 24–35, May–Jun. 2018.
[13] G. Kleppmann, Designing Data-Intensive Applications. Sebastopol, CA, USA: O’Reilly Media, 2017.
[14] S. Sakr, A. Liu, and M. Fayoumi, “The family of MapReduce and large-scale data processing systems,” ACM Computing Surveys, vol. 46, no. 1, pp. 1–44, Jul. 2013.
[15] M. Stonebraker et al., “Requirements for science data bases and scientific data management,” in Proc. 4th Conf. Innovative Data Systems Research (CIDR), Asilomar, CA, USA, 2009, pp. 1–12.
Downloads
How to Cite
[1]E. Williams, “Cloud-Native Data Pipelines for Enterprise Analytics”, IJADSMC, vol. 3, no. 1, pp. 01–13, Apr. 2020, doi: 10.67228/30713498/IJADSMC-2020PI9F3R.