Big Data Processing Using Apache Spark: A Performance Study

  • Authors

    • Dr. R. Kartikeyan Assosciate Professor, Anna University, Chennai, India. Author

    DOI:

    https://doi.org/10.67228/30713498/IJADSMC-2021PI0B9N

    Published 01-02-2021

  • Big Data, Apache Spark, Distributed Computing, In-Memory Processing, Hadoop MapReduce, Performance Evaluation, Scalability, Data Analytics, Cluster Computing, Resource Management

    Issue

    Section

    Articles

    How to Cite

    [1]
    K. R, “Big Data Processing Using Apache Spark: A Performance Study”, IJADSMC, vol. 4, no. 1, pp. 01–16, Jan. 2021, doi: 10.67228/30713498/IJADSMC-2021PI0B9N.
  • Abstract

    The rapid growth of big data across domains such as finance, healthcare, and IoT has exposed the limitations of traditional disk-based systems like Hadoop MapReduce, which suffer from high I/O overhead and inflexible batch processing. Apache Spark addresses these challenges with its in-memory, distributed computing model, enabling faster and more efficient data processing. This paper analyzes Spark’s performance in large-scale data processing, focusing on computational efficiency, scalability, memory usage, and fault tolerance. A comparative study with Hadoop MapReduce shows that Spark achieves 3× to 20× speed improvements, especially for iterative and in-memory workloads. Experimental results also highlight the impact of memory configuration, data partitioning, and data locality on performance. Despite challenges like memory pressure and resource management, Spark proves to be a powerful solution for optimizing big data analytics pipelines.

  • References

    [1] Dean, J., & Ghemawat, S. (2004). MapReduce: Simplified data processing on large clusters. Proceedings of the 6th USENIX Symposium on Operating Systems Design and Implementation (OSDI).

    [2] Dean, J., & Ghemawat, S. (2008). MapReduce: Simplified data processing on large clusters. Communications of the ACM, 51(1), 107–113.

    [3] Zaharia, M., Chowdhury, M., Franklin, M. J., Shenker, S., & Stoica, I. (2010). Spark: Cluster computing with working sets. Proceedings of the 2nd USENIX Workshop on Hot Topics in Cloud Computing (HotCloud).

    [4] Zaharia, M., Xin, R. S., Wendell, P., et al. (2016). Apache Spark: A unified engine for big data processing. Communications of the ACM, 59(11), 56–65.

    [5] Shi, J., Qiu, Y., Minhas, U. F., et al. (2015). Clash of the titans: MapReduce vs. Spark for large-scale data analytics. Proceedings of the VLDB Endowment, 8(13), 2110–2121.

    [6] Lin, J., & Dyer, C. (2010). Data-intensive text processing with MapReduce. Synthesis Lectures on Human Language Technologies, Morgan & Claypool.

    [7] Zaharia, M., Das, T., Li, H., et al. (2013). Discretized streams: Fault-tolerant streaming computation at scale. Proceedings of the 24th ACM Symposium on Operating Systems Principles (SOSP).

    [8] Carbone, P., Katsifodimos, A., Ewen, S., et al. (2015). Apache Flink: Stream and batch processing in a single engine. IEEE Data Engineering Bulletin, 38(4), 28–38.

    [9] Katsifodimos, A., Tzoumas, K., & Markl, V. (2016). Flink: Batch and stream processing in a single engine. IEEE ICDE.

    [10] Abramova, V., & Bernardino, J. (2013). NoSQL databases: MongoDB vs Cassandra. Proceedings of the International Conference on Computer Science and Software Engineering.

    [11] Abadi, D. J., et al. (2016). The design and implementation of modern column-oriented database systems. Foundations and Trends in Databases.

    [12] Akil, B., Zhou, Y., & Röhm, U. (2018). On the usability of Hadoop MapReduce, Apache Spark & Apache Flink for data science. arXiv Technical Report.

    [13] Dolev, S., Florissi, P., Gudes, E., Sharma, S., & Singer, I. (2017). A survey on geographically distributed big-data processing using MapReduce. IEEE Communications Surveys & Tutorials.

    [14] Singh, P., Singh, S., Mishra, P. K., & Garg, R. (2019). RDD-Eclat: Approaches to parallelize Eclat algorithm on Spark RDD framework. Future Generation Computer Systems.

  • Downloads