Synthetic Data Generation Strategies for Pipeline Robustness

  • Authors

    • Dennis Ritchie Research Scientist, Bell Labs, United States. Author
    • Dr. Allen Newell Professor, Carnegie Mellon University, United States. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-2024PI4R3T

    Published 03-09-2024

  • Synthetic Data, Pipeline Robustness, Data Augmentation, Generative Models, Adversarial Robustness, Concept Drift, Machine Learning Evaluation, Data Imbalance, Privacy-Preserving ML, ML Fairness

    Issue

    Section

    Articles

    How to Cite

    [1]
    D. Ritchie and A. Newell, “Synthetic Data Generation Strategies for Pipeline Robustness”, IJDEIC, vol. 7, no. 1, pp. 01–14, Mar. 2024, doi: 10.67228/30715717/IJDEIC-2024PI4R3T.
  • Abstract

    Machine learning pipelines are increasingly deployed in high-stakes domains where robustness against data inconsistencies, bias, and distributional shifts is crucial. However, real-world datasets often suffer from limitations such as data scarcity, class imbalance, and privacy constraints. Synthetic data generation has emerged as a promising strategy to overcome these challenges and enhance pipeline robustness. This paper presents a comprehensive survey and analysis of synthetic data generation techniques, classifying them by data modality, generation method, and application purpose. We explore how synthetic data contributes to the resilience of ML pipelines against failure modes such as concept drift, noise, and adversarial attacks. Through empirical case studies across domains, we demonstrate the practical benefits and limitations of integrating synthetic data into training and evaluation pipelines. Finally, we discuss ethical considerations and outline future research directions toward building more robust, fair, and privacy-preserving ML systems using synthetic data.

  • References

    [1] Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. Advances in Neural Information Processing Systems, 27, 2672–2680.

    [2] Kingma, D. P., & Welling, M. (2014). Auto-Encoding Variational Bayes. arXiv preprint arXiv:1312.6114.

    [3] Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33, 6840–6851.

    [4] Frid-Adar, M., Klang, E., Amitai, M., Goldberger, J., & Greenspan, H. (2018). Synthetic data augmentation using GAN for improved liver lesion classification. IEEE Transactions on Medical Imaging, 38(3), 677–685.

    [5] Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321–357.

    [6] Xu, L., Skoularidou, M., Cuesta-Infante, A., & Veeramachaneni, K. (2019). Modeling tabular data using conditional GAN. Advances in Neural Information Processing Systems, 32, 7335–7345.

    [7] Gajula, S. (2023). A review of anomaly identification in finance frauds using machine learning system. International Journal of Current Engineering and Technology, 13(6), 568–575. https://ijcet.evegenis.org/index.php/ijcet/article/view/820

    [8] Wang, Z., Bovik, A. C., Sheikh, H. R., & Simoncelli, E. P. (2004). Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4), 600–612.

    [9] Jordon, J., Yoon, J., & van der Schaar, M. (2019). PATE-GAN: generating synthetic data with privacy guarantees. International Conference on Learning Representations.

    [10] Sandfort, V., Yan, K., Pickhardt, P. J., & Summers, R. M. (2019). Data augmentation using generative adversarial networks (CycleGAN) to improve generalizability in CT segmentation tasks. Scientific Reports, 9(1), 16884.

    [11] Zhang, H., Cisse, M., Dauphin, Y. N., & Lopez-Paz, D. (2019). Mixup: Beyond empirical risk minimization. International Conference on Learning Representations.

    [12] Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhutdinov, R., Zemel, R., & Bengio, Y. (2015). Show, attend and tell: Neural image caption generation with visual attention. International Conference on Machine Learning.

    [13] Kairouz, P., McMahan, H. B., et al. (2019). Advances and open problems in federated learning. arXiv preprint arXiv:1912.04977.

    [14] Oord, A. v. d., Li, Y., & Vinyals, O. (2018). Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748.

    [15] Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). "Why should I trust you?": Explaining the predictions of any classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1135–1144.

    [16] Liu, X., Wang, H., & Gao, J. (2021). Diffusion models for generative data augmentation in natural language processing. arXiv preprint arXiv:2107.08924.

  • Downloads