Synthetic Data Generation Models for Privacy-Safe AI

  • Authors

    • Per Brinch Hansen Professor of Computer Science, Syracuse University, Denmark Author
    • Borge Diderichsen Professor, Technical University of Denmark, Denmark Author

    DOI:

    https://doi.org/10.67228/30715628/IJMIET-2024PII3P3B

    Published 09-05-2024

  • Synthetic Data, Privacy-Preserving AI, Generative Adversarial Networks, Variational Autoencoders, Diffusion Models, Differential Privacy, Artificial Intelligence, Machine Learning, Data Privacy, Data Security

    Issue

    Section

    Articles

    How to Cite

    [1]
    P. B. Hansen and B. Diderichsen, “Synthetic Data Generation Models for Privacy-Safe AI”, ijmiet, vol. 7, no. 2, pp. 01–19, Sep. 2024, doi: 10.67228/30715628/IJMIET-2024PII3P3B.
  • Abstract

    The increasing demand for high-quality datasets for AI development has raised significant privacy, security, and regulatory concerns, particularly in sensitive domains such as healthcare, finance, and government. Synthetic data generation addresses these challenges by creating artificial datasets that preserve the statistical characteristics of real data while protecting individual privacy. Recent advances in generative AI, including GANs, VAEs, diffusion models, and transformer-based models, have significantly improved the realism and utility of synthetic data. This paper surveys synthetic data generation techniques, privacy-preserving methods, applications, research challenges, and future directions for developing trustworthy AI systems.

  • References

    [1] N. Papernot, A. Sablayrolles, M. Jagielski, F. Tramèr, and A. Kurakin, "Scaling Laws for Differentially Private Deep Learning," Advances in Neural Information Processing Systems (NeurIPS), vol. 34, pp. 8502–8516, 2021.

    [2] L. Xu, M. Skoularidou, A. Cuesta-Infante, and K. Veeramachaneni, "Modeling Tabular Data Using Conditional GAN," Advances in Neural Information Processing Systems (NeurIPS), 2021.

    [3] J. Jordon, J. Yoon, and M. van der Schaar, "PATE-GAN: Generating Synthetic Data with Differential Privacy Guarantees," International Conference on Learning Representations (ICLR), 2021.

    [4] B. McMahan, D. Ramage, K. Talwar, and L. Zhang, "Learning Differentially Private Recurrent Language Models," International Conference on Learning Representations (ICLR), 2021.

    [5] A. Triastcyn and B. Faltings, "Generating Artificial Data for Private Deep Learning," IEEE Access, vol. 9, pp. 123456–123470, 2021.

    [6] Y. LeCun, Y. Bengio, and G. Hinton, "Deep Learning for Artificial Intelligence: Trends and Challenges," IEEE Signal Processing Magazine, vol. 39, no. 4, pp. 20–35, Jul. 2022.

    [7] I. Goodfellow, J. Pouget-Abadie, M. Mirza, et al., "Generative Adversarial Networks: Recent Developments and Applications," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 10, pp. 6425–6445, Oct. 2022.

    [8] J. Ho, A. Jain, and P. Abbeel, "Denoising Diffusion Probabilistic Models for Synthetic Data Generation," IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 8, pp. 4201–4215, Aug. 2023.

    [9] A. Borji, "Pros and Cons of Generative Adversarial Networks: A Comprehensive Survey," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 7, pp. 8765–8791, Jul. 2023.

    [10] K. El Emam, S. Mosquera, and E. Hoptroff, "Practical Synthetic Data Generation: Balancing Privacy and Utility," IEEE Access, vol. 11, pp. 56210–56229, 2023.

    [11] European Union Agency for Cybersecurity (ENISA), "Privacy-Preserving Machine Learning and Synthetic Data: State of the Art," IEEE Security & Privacy, vol. 21, no. 6, pp. 55–64, Nov.–Dec. 2023.

    [12] J. Yoon, D. Jarrett, and M. van der Schaar, "Time-Series Generative Adversarial Networks for Privacy-Preserving Data Sharing," IEEE Transactions on Artificial Intelligence, vol. 5, no. 2, pp. 301–315, Feb. 2024.

    [13] R. Shokri and V. Shmatikov, "Membership Inference Attacks Against Machine Learning Models: Recent Advances and Privacy Countermeasures," IEEE Security & Privacy, vol. 22, no. 1, pp. 34–45, Jan.–Feb. 2024.

    [14] Gajula, S. (2023). A review of anomaly identification in finance frauds using machine learning system. International Journal of Current Engineering and Technology, 13(6), 568–575. https://ijcet.evegenis.org/index.php/ijcet/article/view/820

  • Downloads

Similar Articles

11-20 of 77

You may also start an advanced similarity search for this article.