Data-Centric Security for Generative AI Systems: Protecting Sensitive Data against Leakage, Inference Attacks, and Unauthorized Model Access

  • Authors

    • Santosh Kumar Jadala Cyber Security & Business Analysis Specialist, Independent Researcher, USA. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-V9I3P101

    Published 08-02-2026

  • Generative Artificial Intelligence, Data-Centric Security, Data Leakage, Inference Attacks, Model Security, Sensitive Data Protection, Access Control, Privacy-Preserving AI

    Issue

    Section

    Articles

    How to Cite

    [1]
    S. K. Jadala, “Data-Centric Security for Generative AI Systems: Protecting Sensitive Data against Leakage, Inference Attacks, and Unauthorized Model Access”, IJDEIC, vol. 9, no. 3, pp. 01–32, Aug. 2026, doi: 10.67228/30715717/IJDEIC-V9I3P101.
  • Abstract

    The growing use of generative artificial intelligence in enterprise and public-sector environments has introduced significant concerns regarding the protection of sensitive data. Generative AI systems process large volumes of information across training, fine-tuning, retrieval, inference, and output stages, creating multiple opportunities for unauthorized disclosure and misuse. This study examines data-centric security as an approach for protecting sensitive information against data leakage, membership inference, model inversion, prompt-based attacks, retrieval-related exposure, model extraction, and unauthorized access to AI services. It reviews major attack surfaces across the generative AI lifecycle and evaluates security controls including data classification, minimization, encryption, privacy-preserving learning, identity and access management, secure retrieval-augmented generation, output filtering, data loss prevention, and continuous monitoring. Based on these findings, the study proposes a data-centric security framework that links data sensitivity, access policies, model controls, and governance requirements throughout the AI lifecycle. The framework emphasizes that security controls should remain tied to sensitive information regardless of where the data is stored, processed, retrieved, or generated. The study provides practical guidance for organizations seeking to reduce privacy and confidentiality risks while maintaining controlled and accountable use of generative AI systems.

  • References

    [1] Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., & Zhang, L. (2016). Deep learning with differential privacy. Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, 308–318. doi:10.1145/2976749.2978318.

    [2] Aerni, M., Rando, J., Debenedetti, E., Carlini, N., Ippolito, D., & Tramèr, F. (2025). Measuring non-adversarial reproduction of training data in large language models. International Conference on Learning Representations (ICLR 2025).

    [3] An, B., Zhang, S., & Dredze, M. (2025). RAG LLMs are not safer: A safety analysis of retrieval-augmented generation for large language models. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, 5444–5474. doi: 10.18653/v1/2025.naacl-long.281.

    [4] Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., & Song, D. (2019). The Secret Sharer: Evaluating and testing unintended memorization in neural networks. 28th USENIX Security Symposium (USENIX Security 19), 267–284.

    [5] Carlini, N., Tramèr, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, Ú., Oprea, A., & Raffel, C. (2021). Extracting training data from large language models. 30th USENIX Security Symposium (USENIX Security 21), 2633–2650.

    [6] Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., & Zhang, C. (2023). Quantifying memorization across neural language models. International Conference on Learning Representations (ICLR 2023).

    [7] Carlini, N., Hayes, J., Nasr, M., Jagielski, M., Sehwag, V., Tramèr, F., Balle, B., Ippolito, D., & Wallace, E. (2023). Extracting training data from diffusion models. 32nd USENIX Security Symposium (USENIX Security 23), 5253–5270.

    [8] Carlini, N., Paleka, D., Dvijotham, K. D., Steinke, T., Hayase, J., Cooper, A. F., Lee, K., Jagielski, M., Nasr, M., Conmy, A., Wallace, E., Rolnick, D., & Tramèr, F. (2024). Stealing part of a production language model. Proceedings of the 41st International Conference on Machine Learning, 235, 5680–5705.

    [9] Kunaparaju, C. (2025). AI-Driven Cyber Defense Systems: Strengthening National Security through Intelligent Threat Prediction and Response. Algora, 2(1), 1-30.

    [10] Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G. J., Tramèr, F., Hassani, H., & Wong, E. (2024). JailbreakBench: An open robustness benchmark for jailbreaking large language models. Advances in Neural Information Processing Systems, 37. doi:10.52202/079017-1745.

    [11] Potla, R. B. (2024). Optimizing extended warehouse management for make-to-order plants: Slotting, wave picking, and yard orchestration at scale. Journal of Computer Science and Technology Studies, 6(3), 181-192.

    [12] Cheng, Y., Zhang, L., Wang, J., Yuan, M., & Yao, Y. (2025). RemoteRAG: A privacy-preserving LLM cloud RAG service. Findings of the Association for Computational Linguistics: ACL 2025, 3820–3837. doi: 10.18653/v1/2025.findings-acl.197.

    [13] Flemings, J., Jiang, B., Zhang, W., Takhirov, Z., & Annavaram, M. (2025). Estimating privacy leakage of augmented contextual knowledge in language models. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 25092–25108. doi: 10.18653/v1/2025.acl-long.1220.

    [14] Fredrikson, M., Jha, S., & Ristenpart, T. (2015). Model inversion attacks that exploit confidence information and basic countermeasures. Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security, 1322–1333. doi:10.1145/2810103.2813677.

    [15] Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security. doi:10.1145/3605764.3623985.

    [16] Huang, J., Shao, H., & Chang, K. C.-C. (2022). Are large-pre-trained language models leaking your personal information? Findings of the Association for Computational Linguistics: EMNLP 2022, 2038–2047. doi: 10.18653/v1/2022.findings-emnlp.148.

    [17] Huang, Y., Gupta, S., Zhong, Z., Li, K., & Chen, D. (2023). Privacy implications of retrieval-based language models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 14887–14902.

    [18] Potla, R. B. (2024). A SOX/ITAR-aligned global ERP template for multi-plant manufacturers: Governance patterns and controls. J Artif Intell Mach Learn & Data Sci, 2(2), 3222-3232.

    [19] Begimher, D., Leo, C., Huang, J., Gaw, P., & Zheng, B. (2026). SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents. arXiv preprint arXiv:2604.12040.

    [20] Kandpal, N., Wallace, E., & Raffel, C. (2022). Deduplicating training data mitigates privacy risks in language models. Proceedings of the 39th International Conference on Machine Learning, 162, 10697–10707.

    [21] Li, X., Tramèr, F., Liang, P., & Hashimoto, T. (2022). Large language models can be strong differentially private learners. International Conference on Learning Representations (ICLR 2022).

    [22] Lukas, N., Salem, A., Sim, R., Tople, S., Wutschitz, L., & Zanella-Béguelin, S. (2023). Analyzing leakage of personally identifiable information in language models. 2023 IEEE Symposium on Security and Privacy, 346–363.

    [23] Nasr, M., Shokri, R., & Houmansadr, A. (2019). Comprehensive privacy analysis of deep learning: Passive and active white-box inference attacks against centralized and federated learning. 2019 IEEE Symposium on Security and Privacy, 739–753.

    [24] Nasr, M., Rando, J., Carlini, N., Hayase, J., Jagielski, M., Cooper, A. F., Ippolito, D., Choquette-Choo, C. A., Tramèr, F., & Lee, K. (2025). Scalable extraction of training data from aligned, production language models. International Conference on Learning Representations (ICLR 2025).

    [25] Perez, F., & Ribeiro, I. (2022). Ignore previous prompt: Attack techniques for language models. arXiv preprint arXiv:2211.09527.

    [26] Shokri, R., Stronati, M., Song, C., & Shmatikov, V. (2017). Membership inference attacks against machine learning models. 2017 IEEE Symposium on Security and Privacy, 3–18. doi:10.1109/SP.2017.41.

    [27] Kunaparaju, C. (2024). The Role of Artificial Intelligence in Safeguarding Critical National Infrastructure against Cyberattacks. Journal of Electrical Systems, 20, 5403-5419.

    [28] Somepalli, G., Singla, V., Goldblum, M., Geiping, J., & Goldstein, T. (2023). Diffusion art or digital forgery? Investigating data replication in diffusion models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

    [29] Potla, R. (2023). Designing a BTP-centric integration mesh for shop-floor IoT, MES and ERP in discrete manufacturing. Journal of Artificial Intelligence, Machine Learning and Data Science, 1(2), 1-8.

    [30] Tramèr, F., Zhang, F., Juels, A., Reiter, M. K., & Ristenpart, T. (2016). Stealing machine learning models via prediction APIs. 25th USENIX Security Symposium (USENIX Security 16), 601–618.

    [31] Wang, B., Chen, W., Pei, H., Xie, C., Kang, M., Zhang, C., Xu, C., Xiong, Z., Dutta, R., Schaeffer, R., Truong, S., Arora, S., Mazeika, M., Hendrycks, D., Lin, Z., Cheng, Y., Koyejo, S., Song, D., & Li, B. (2023). DecodingTrust: A comprehensive assessment of trustworthiness in GPT models. Advances in Neural Information Processing Systems, 36.

    [32] Wang, B., He, W., Zeng, S., Xiang, Z., Xing, Y., Tang, J., & He, P. (2025). Unveiling privacy risks in LLM agent memory. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 25241–25260. doi: 10.18653/v1/2025.acl-long.1227.

    [33] Leo, C., Dykyi, A., Cortegaca, D., Begimher, D., & Jha, P. (2026). ThreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping. arXiv preprint arXiv:2607.27528.

    [34] Wei, A., Haghtalab, N., & Steinhardt, J. (2023). Jailbroken: How does LLM safety training fail? Advances in Neural Information Processing Systems, 36.

    [35] Yi, J., et al. (2025). Benchmarking and defending against indirect prompt injection attacks on large language models. Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining. doi:10.1145/3690624.3709179.

    [36] Zeng, S., Zhang, J., He, P., Xing, Y., Liu, Y., Xu, H., Ren, J., Wang, S., Yin, D., Chang, Y., & Tang, J. (2024). The good and the bad: Exploring privacy issues in retrieval-augmented generation (RAG). Findings of the Association for Computational Linguistics: ACL 2024, 4505–4524. doi: 10.18653/v1/2024.findings-acl.267.

    [37] Zou, W., Geng, R., Wang, B., & Jia, J. (2025). PoisonedRAG: Knowledge corruption attacks to retrieval-augmented generation of large language models. 34th USENIX Security Symposium (USENIX Security 25), 3827–3844.

  • Downloads