AI-Augmented Data Quality Monitoring in Real-Time Data Pipelines on AWS

  • Authors

    • Michael Anderson Director of Information Technology, ABC Technologies Inc., USA. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-2018PI8R2N

    Published 06-11-2018

  • Data Quality Monitoring, Real-Time Data Pipelines, Artificial Intelligence, AWS Cloud, Machine Learning, Anomaly Detection, Amazon Kinesis, SageMaker, Data Drift, Data Observability, Streaming Analytics, Cloud-Native Architecture

    Issue

    Section

    Articles

    How to Cite

    [1]
    M. Anderson, “AI-Augmented Data Quality Monitoring in Real-Time Data Pipelines on AWS”, IJDEIC, vol. 1, no. 1, pp. 01–14, Jun. 2018, doi: 10.67228/30715717/IJDEIC-2018PI8R2N.
  • Abstract

    Ensuring high data quality is essential for the success of real-time data analytics, particularly in large-scale cloud environments such as AWS. Traditional rule-based monitoring approaches are often brittle, labor-intensive, and ill-suited for dynamic data patterns. This paper proposes an AI-augmented approach to data quality monitoring within real-time data pipelines on AWS. We present a reference architecture that integrates machine learning models for anomaly detection, data drift analysis, and predictive quality assessment into the AWS streaming ecosystem. The pipeline utilizes services such as Amazon Kinesis, AWS Lambda, SageMaker, and CloudWatch for scalable, low-latency data processing and observability. Through empirical evaluation, we demonstrate the effectiveness of AI-enhanced monitoring in identifying and mitigating quality issues with minimal human intervention, highlighting improvements in precision, recall, and operational efficiency. This approach not only improves trust in data-driven decisions but also offers a scalable solution for modern data engineering practices.

  • References

    [1] Wang, R. Y., & Strong, D. M. (1996). Beyond accuracy: What data quality means to data consumers. Journal of Management Information Systems, 12(4), 5–33. https://doi.org/10.1080/07421222.1996.11518099

    [2] Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3), Article 15. https://doi.org/10.1145/1541880.1541882

    [3] Zaharia, M., Xin, R. S., Wendell, P., Das, T., Armbrust, M., Dave, A., Meng, X., Rosen, J., Venkataraman, S., Franklin, M. J., Ghodsi, A., Gonzalez, J., Shenker, S., & Stoica, I. (2016). Apache Spark: A unified engine for big data processing. Communications of the ACM, 59(11), 56–65. https://doi.org/10.1145/2934664

    [4] Kreps, J., Narkhede, N., & Rao, J. (2011). Kafka: A distributed messaging system for log processing. Proceedings of the NetDB Workshop, 1–7.

    [5] Kingma, D. P., & Welling, M. (2014). Auto-encoding variational Bayes. Proceedings of the International Conference on Learning Representations (ICLR).

    [6] Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys, 46(4), Article 44. https://doi.org/10.1145/2523813

    [7] Kimball, R., & Caserta, J. (2004). The data warehouse ETL toolkit: Practical techniques for extracting, cleaning, conforming, and delivering data. Wiley.

    [8] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). https://doi.org/10.1145/2939672.2939785

    [9] Batini, C., & Scannapieco, M. (2016). Data and information quality: Dimensions, principles and techniques. Springer. https://doi.org/10.1007/978-3-319-24106-7

    [10] Jonas, E., Schleier-Smith, J., Sreekanti, V., Tsai, C.-C., Khandelwal, A., Pu, Q., Shankar, V., Carreira, J., Krauth, K., Yadwadkar, N., Gonzalez, J., Popa, R., Stoica, I., & Patterson, D. (2017). Cloud programming simplified: A Berkeley view on serverless computing. arXiv Preprint arXiv:1902.03383.

    [11] Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Le, Q. V., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., Ng, A. Y., & others. (2012). Large scale distributed deep networks. Advances in Neural Information Processing Systems, 25, 1223–1231.

    [12] Kifer, D., & Machanavajjhala, A. (2011). No free lunch in data privacy. In Proceedings of the ACM SIGMOD International Conference on Management of Data (pp. 193–204). https://doi.org/10.1145/1989323.1989345

    [13] Jaggi, M., Smith, V., Takáč, M., Terhorst, J., Krishnan, S., Hofmann, T., & Jordan, M. I. (2014). Communication-efficient learning of deep networks from decentralized data. In Proceedings of the 17th International Conference on Artificial Intelligence and Statistics (AISTATS) (pp. 354–362).

    [14] Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press. https://www.deeplearningbook.org

    [15] Nair, V., & Hinton, G. E. (2010). Rectified linear units improve restricted Boltzmann machines. In Proceedings of the 27th International Conference on Machine Learning (pp. 807–814).

  • Downloads