Automated Data Quality Monitoring in Streaming Data Transformations
-
DOI:
https://doi.org/10.67228/30715717/IJDEIC-2018PII2F6APublished 07-06-2018
Streaming Data, Data Quality Monitoring, Real-Time Data Pipelines, Apache Kafka, Anomaly Detection, Schema Drift, Spark Streaming, Data Observability, Automated Validation, Streaming Analytics Issue
Section
ArticlesHow to Cite
[1]K. Taylor and A. Reza, “Automated Data Quality Monitoring in Streaming Data Transformations”, IJDEIC, vol. 1, no. 2, pp. 01–12, Jul. 2018, doi: 10.67228/30715717/IJDEIC-2018PII2F6A.Abstract
As real-time data becomes integral to decision-making, ensuring the quality of streaming data is critical to maintaining system reliability and business value. This paper explores automated data quality monitoring techniques specifically tailored for streaming data transformations. It investigates common data quality issues such as latency, duplication, schema drift, null values, and data inconsistency in real-time pipelines. The study presents a framework for integrating automated monitoring tools within streaming architectures using technologies like Apache Kafka, Apache Fink, and Spark Streaming. Machine learning-based anomaly detection models and rule-based engines are examined for detecting quality issues in-flight. Through experimental evaluation, the paper demonstrates how automated monitoring improves data observability, reduces downtime, and supports proactive data governance. The findings provide a roadmap for organizations seeking to implement scalable, real-time data quality solutions within their streaming ecosystems.
References
[1] Wang, R. Y., & Strong, D. M. (1996). Beyond accuracy: What data quality means to data consumers. Journal of Management Information Systems, 12(4), 5–33. https://doi.org/10.1080/07421222.1996.11518099
[2] Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J. F., & Dennison, D. (2015). Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems, 28, 2503–2511.
[3] Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3), Article 15. https://doi.org/10.1145/1541880.1541882
[4] Zaharia, M., Das, T., Li, H., Hunter, T., Shenker, S., & Stoica, I. (2013). Discretized streams: Fault-tolerant streaming computation at scale. In Proceedings of the Twenty-Fourth ACM Symposium on Operating Systems Principles (pp. 423–438). https://doi.org/10.1145/2517349.2522737
[5] Batini, C., & Scannapieco, M. (2016). Data and information quality: Dimensions, principles and techniques. Springer. https://doi.org/10.1007/978-3-319-24106-7
[6] Kimball, R., & Ross, M. (2013). The data warehouse toolkit: The definitive guide to dimensional modeling (3rd ed.). Wiley.
[7] Cugola, G., & Margara, A. (2012). Processing flows of information: From data stream to complex event processing. ACM Computing Surveys, 44(3), 1–62. https://doi.org/10.1145/2187671.2187677
[8] Menascé, D. A., & Almeida, V. A. F. (2002). Capacity planning for Web services: Metrics, models, and methods. Prentice Hall.
[9] Aggarwal, C. C. (2013). Outlier analysis (2nd ed.). Springer. https://doi.org/10.1007/978-1-4614-6396-2
[10] English, L. P. (2009). Information quality applied: Best practices for improving business information, processes and systems. Wiley.
[11] Akidau, T., Bradshaw, R., Chambers, C., Chernyak, S., Fernández-Moctezuma, R. J., Lax, R., McVeety, S., Mills, D., Perry, F., Schmidt, E., & Whittle, S. (2015). The dataflow model: A practical approach to balancing correctness, latency, and cost in massive-scale, unbounded, out-of-order data processing. Proceedings of the VLDB Endowment, 8(12), 1792–1803. https://doi.org/10.14778/2824032.2824076
[12] Kreps, J., Narkhede, N., & Rao, J. (2011). Kafka: A distributed messaging system for log processing. Proceedings of the NetDB Workshop, 1–7.
[13] Ahmed, M., Mahmood, A. N., & Hu, J. (2016). A survey of network anomaly detection techniques. Journal of Network and Computer Applications, 60, 19–31. https://doi.org/10.1016/j.jnca.2015.11.016
[14] Dunning, T., & Friedman, E. (2014). Time series databases: New ways to store and access data. O’Reilly Media.
[15] Dean, J., & Ghemawat, S. (2008). MapReduce: Simplified data processing on large clusters. Communications of the ACM, 51(1), 107–113. https://doi.org/10.1145/1327452.1327492
Downloads
How to Cite
[1]K. Taylor and A. Reza, “Automated Data Quality Monitoring in Streaming Data Transformations”, IJDEIC, vol. 1, no. 2, pp. 01–12, Jul. 2018, doi: 10.67228/30715717/IJDEIC-2018PII2F6A.