Big Data Processing Using Apache Spark: A Performance Study
-
DOI:
https://doi.org/10.67228/30713498/IJADSMC-2021PI0B9NPublished 01-02-2021
Big Data, Apache Spark, Distributed Computing, In-Memory Processing, Hadoop MapReduce, Performance Evaluation, Scalability, Data Analytics, Cluster Computing, Resource Management Issue
Section
ArticlesHow to Cite
[1]K. R, “Big Data Processing Using Apache Spark: A Performance Study”, IJADSMC, vol. 4, no. 1, pp. 01–16, Jan. 2021, doi: 10.67228/30713498/IJADSMC-2021PI0B9N.Abstract
The rapid growth of big data across domains such as finance, healthcare, and IoT has exposed the limitations of traditional disk-based systems like Hadoop MapReduce, which suffer from high I/O overhead and inflexible batch processing. Apache Spark addresses these challenges with its in-memory, distributed computing model, enabling faster and more efficient data processing. This paper analyzes Spark’s performance in large-scale data processing, focusing on computational efficiency, scalability, memory usage, and fault tolerance. A comparative study with Hadoop MapReduce shows that Spark achieves 3× to 20× speed improvements, especially for iterative and in-memory workloads. Experimental results also highlight the impact of memory configuration, data partitioning, and data locality on performance. Despite challenges like memory pressure and resource management, Spark proves to be a powerful solution for optimizing big data analytics pipelines.
References
[1] Dean, J., & Ghemawat, S. (2004). MapReduce: Simplified data processing on large clusters. Proceedings of the 6th USENIX Symposium on Operating Systems Design and Implementation (OSDI).
[2] Dean, J., & Ghemawat, S. (2008). MapReduce: Simplified data processing on large clusters. Communications of the ACM, 51(1), 107–113.
[3] Zaharia, M., Chowdhury, M., Franklin, M. J., Shenker, S., & Stoica, I. (2010). Spark: Cluster computing with working sets. Proceedings of the 2nd USENIX Workshop on Hot Topics in Cloud Computing (HotCloud).
[4] Zaharia, M., Xin, R. S., Wendell, P., et al. (2016). Apache Spark: A unified engine for big data processing. Communications of the ACM, 59(11), 56–65.
[5] Shi, J., Qiu, Y., Minhas, U. F., et al. (2015). Clash of the titans: MapReduce vs. Spark for large-scale data analytics. Proceedings of the VLDB Endowment, 8(13), 2110–2121.
[6] Lin, J., & Dyer, C. (2010). Data-intensive text processing with MapReduce. Synthesis Lectures on Human Language Technologies, Morgan & Claypool.
[7] Zaharia, M., Das, T., Li, H., et al. (2013). Discretized streams: Fault-tolerant streaming computation at scale. Proceedings of the 24th ACM Symposium on Operating Systems Principles (SOSP).
[8] Carbone, P., Katsifodimos, A., Ewen, S., et al. (2015). Apache Flink: Stream and batch processing in a single engine. IEEE Data Engineering Bulletin, 38(4), 28–38.
[9] Katsifodimos, A., Tzoumas, K., & Markl, V. (2016). Flink: Batch and stream processing in a single engine. IEEE ICDE.
[10] Abramova, V., & Bernardino, J. (2013). NoSQL databases: MongoDB vs Cassandra. Proceedings of the International Conference on Computer Science and Software Engineering.
[11] Abadi, D. J., et al. (2016). The design and implementation of modern column-oriented database systems. Foundations and Trends in Databases.
[12] Akil, B., Zhou, Y., & Röhm, U. (2018). On the usability of Hadoop MapReduce, Apache Spark & Apache Flink for data science. arXiv Technical Report.
[13] Dolev, S., Florissi, P., Gudes, E., Sharma, S., & Singer, I. (2017). A survey on geographically distributed big-data processing using MapReduce. IEEE Communications Surveys & Tutorials.
[14] Singh, P., Singh, S., Mishra, P. K., & Garg, R. (2019). RDD-Eclat: Approaches to parallelize Eclat algorithm on Spark RDD framework. Future Generation Computer Systems.
Downloads
How to Cite
[1]K. R, “Big Data Processing Using Apache Spark: A Performance Study”, IJADSMC, vol. 4, no. 1, pp. 01–16, Jan. 2021, doi: 10.67228/30713498/IJADSMC-2021PI0B9N.