Scalable Event-Driven Architectures for Real-Time Big Data Applications

  • Authors

    • Dr. Nandhini Ravi Assistant Professor, VIT University, India. Author

    DOI:

    https://doi.org/10.67228/30713498/IJADSMC-2024PII9F8M

    Published 09-03-2024

  • Event-Driven Architecture, Big Data Analytics, Real-Time Processing, Stream Analytics, Apache Kafka, Apache Flink, Microservices, Scalability, Cloud Computing, Distributed Systems

    Issue

    Section

    Articles

    How to Cite

    [1]
    N. Ravi, “Scalable Event-Driven Architectures for Real-Time Big Data Applications”, IJADSMC, vol. 7, no. 2, pp. 01–17, Sep. 2024, doi: 10.67228/30713498/IJADSMC-2024PII9F8M.
  • Abstract

    Scalable Event-Driven Architectures (EDAs) provide an efficient approach for real-time big data processing by enabling low-latency, asynchronous, and continuous handling of high-speed data streams generated from IoT devices, cloud platforms, social media, and enterprise systems. This study presents a scalable framework that integrates distributed event brokers, stream processing engines, cloud-native microservices, and scalable storage solutions. The architecture utilizes publish-subscribe communication, event sourcing, and distributed stream analytics to support real-time decision-making and dynamic scalability. Technologies such as Apache Kafka, Apache Flink, Apache Spark Streaming, and Kubernetes enhance throughput, fault tolerance, and system resilience. Key components include event producers, brokers, stream processors, consumers, and monitoring services. Performance evaluation based on throughput, latency, scalability, fault tolerance, and resource utilization demonstrates that EDAs outperform traditional batch-processing and request-response systems. The framework also addresses challenges such as event ordering, state management, event replay, and observability through checkpointing, event partitioning, distributed tracing, and container orchestration. The proposed architecture is applicable to financial analytics, smart cities, healthcare, e-commerce, cybersecurity, and industrial automation. Overall, EDAs offer a scalable, reliable, and flexible foundation for next-generation real-time analytics and big data applications.

  • References

    [1] T. Akidau, S. Chernyak, and R. Lax, Streaming Systems: The What, Where, When, and How of Large-Scale Data Processing. Sebastopol, CA, USA: O'Reilly Media, 2018.

    [2] J. Kreps, N. Narkhede, and J. Rao, “Kafka: A Distributed Messaging System for Log Processing,” in Proc. NetDB Workshop, Athens, Greece, 2011, pp. 1–7.

    [3] P. Carbone, G. Katsifodimos, S. Ewen, V. Markl, S. Haridi, and K. Tzoumas, “Apache Flink: Stream and Batch Processing in a Single Engine,” IEEE Data Engineering Bulletin, vol. 38, no. 4, pp. 28–38, 2015.

    [4] M. Zaharia, T. Das, H. Li, T. Hunter, S. Shenker, and I. Stoica, “Discretized Streams: Fault-Tolerant Streaming Computation at Scale,” in Proc. ACM Symposium on Operating Systems Principles (SOSP), Farmington, PA, USA, 2013, pp. 423–438.

    [5] N. Marz and J. Warren, Big Data: Principles and Best Practices of Scalable Real-Time Data Systems. Greenwich, CT, USA: Manning Publications, 2015.

    [6] [6] T. White, Hadoop: The Definitive Guide, 4th ed. Sebastopol, CA, USA: O’Reilly Media, 2015.

    [7] G. Hohpe and B. Woolf, Enterprise Integration Patterns: Designing, Building, and Deploying Messaging Solutions. Boston, MA, USA: Addison-Wesley, 2004.

    [8] M. Fowler, “What Do You Mean by Event-Driven Architecture?” IEEE Software, vol. 34, no. 3, pp. 15–17, 2017.

    [9] B. Gedik, S. Schneider, M. Hirzel, and K. Wu, “Elastic Scaling for Data Stream Processing,” IEEE Transactions on Parallel and Distributed Systems, vol. 25, no. 6, pp. 1447–1463, Jun. 2014.

    [10] S. Newman, Building Microservices: Designing Fine-Grained Systems. Sebastopol, CA, USA: O’Reilly Media, 2021.

    [11] B. Burns, B. Grant, D. Oppenheimer, E. Brewer, and J. Wilkes, “Borg, Omega, and Kubernetes,” Communications of the ACM, vol. 59, no. 5, pp. 50–57, May 2016.

    [12] N. Dragoni, S. Dustdar, S. Larsen, and M. Mazzara, “Microservices: Migration of a Mission Critical System,” IEEE Software, vol. 35, no. 3, pp. 56–63, 2018.

    [13] M. Kleppmann, Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems. Sebastopol, CA, USA: O’Reilly Media, 2017.

    [14] S. Chandramouli, J. Goldstein, and D. Maier, “On the Performance of Distributed Stream Processing Engines,” Proceedings of the VLDB Endowment, vol. 8, no. 12, pp. 1714–1725, 2015.

    [15] V. K. Vavilapalli et al., “Apache Hadoop YARN: Yet Another Resource Negotiator,” in Proc. ACM Symposium on Cloud Computing (SoCC), San Jose, CA, USA, 2013, pp. 1–16.

    [16] Gajula, S. (2023). A review of anomaly identification in finance frauds using machine learning system. International Journal of Current Engineering and Technology, 13(6), 568–575. https://ijcet.evegenis.org/index.php/ijcet/article/view/820.

  • Downloads