Securing Retrieval-Augmented Generation Systems: Threat Modeling, Data Poisoning, Prompt Injection, and Access Control for Enterprise AI

  • Authors

    • Santosh Kumar Jadala Cyber Security & Business Analysis Specialist, Independent Researcher, USA. Author

    DOI:

    https://doi.org/10.67228/30713498/IJADSMC-V9I3P101

    Published 07-04-2026

  • Retrieval-Augmented Generation, RAG Security, Prompt Injection, Data Poisoning, Access Control, Enterprise AI, Zero Trust, Large Language Models

    Issue

    Section

    Articles

    How to Cite

    [1]
    S. K. Jadala, “Securing Retrieval-Augmented Generation Systems: Threat Modeling, Data Poisoning, Prompt Injection, and Access Control for Enterprise AI”, IJADSMC, vol. 9, no. 3, pp. 01–14, Jul. 2026, doi: 10.67228/30713498/IJADSMC-V9I3P101.
  • Abstract

    Retrieval-Augmented Generation (RAG) has emerged as an important architecture for enhancing large language models with external enterprise knowledge, enabling more accurate, contextual, and domain-specific responses. However, integrating retrieval mechanisms with generative models introduces a distinct security attack surface involving untrusted data sources, vector databases, retrieval pipelines, prompts, and access-control mechanisms. This study examines the security risks affecting enterprise RAG systems, with particular attention to data poisoning, prompt injection, unauthorized knowledge retrieval, information leakage, and inadequate access control. It develops a threat-oriented security framework that maps adversarial activities across the RAG lifecycle, from data ingestion and embedding generation to retrieval, prompt construction, and response generation. The proposed approach combines secure data provenance, poisoning detection, prompt isolation, retrieval-level authorization, role-based and attribute-based access controls, Zero Trust principles, continuous monitoring, and auditability. The study further emphasizes defense-in-depth as a necessary strategy for protecting sensitive enterprise knowledge while preserving the operational benefits of RAG. By integrating established cybersecurity principles with emerging large language model threats, the research provides a structured foundation for securing enterprise RAG deployments. The study concludes that effective RAG security requires coordinated controls across data, retrieval, generation, identity, and governance layers.

  • References

    [1] Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474. https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html

    [2] Kunaparaju, C. (2025). AI-Driven Cyber Defense Systems: Strengthening National Security through Intelligent Threat Prediction and Response. Algora, 2(1), 1-30.

    [3] Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W. (2020). Dense passage retrieval for open-domain question answering. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 6769–6781. https://doi.org/10.18653/v1/2020.emnlp-main.550

    [4] Guu, K., Lee, K., Tung, Z., Pasupat, P., & Chang, M. (2020). Retrieval augmented language model pre-training. Proceedings of the 37th International Conference on Machine Learning, 119, 3929–3938. https://proceedings.mlr.press/v119/guu20a.html

    [5] Izacard, G., & Grave, E. (2021). Leveraging passage retrieval with generative models for open domain question answering. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics, 874–880.

    https://doi.org/10.18653/v1/2021.eacl-main.74

    [6] Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec ’23). https://doi.org/10.1145/3605764.3623985

    [7] Schulhoff, S., Pinto, J., Khan, A., Bouchard, L. F., Si, C., Anati, S., Tagliabue, V., Kost, A., Carnahan, C., & Boyd-Graber, J. (2023). Ignore this title and HackAPrompt: Exposing systemic vulnerabilities of LLMs through a global prompt hacking competition. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4945–4977.

    https://doi.org/10.18653/v1/2023.emnlp-main.302

    [8] Yip, D. W., Esmradi, A., & Chan, C. F. (2023). A novel evaluation framework for assessing resilience against prompt injection attacks in large language models. 2023 IEEE Asia-Pacific Conference on Computer Science and Data Engineering (CSDE), 1–5. https://doi.org/10.1109/CSDE59766.2023.10487667

    [9] Wallace, E., Feng, S., Kandpal, N., Gardner, M., & Singh, S. (2019). Universal adversarial triggers for attacking and analyzing NLP. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, 2153–2162. https://doi.org/10.18653/v1/D19-1221

    [10] Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., & Irving, G. (2022). Red teaming language models with language models. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3419–3448. https://doi.org/10.18653/v1/2022.emnlp-main.225

    [11] Zhao, S., Wen, J., Luu, A., Zhao, J., & Fu, J. (2023). Prompt as triggers for backdoor attack: Examining the vulnerability in language models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12303–12317.

    https://doi.org/10.18653/v1/2023.emnlp-main.757

    [12] Li, L., Song, D., Li, X., Zeng, J., Ma, R., & Qiu, X. (2021). Backdoor attacks on pre-trained models by layerwise weight poisoning. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 3023–3032. https://doi.org/10.18653/v1/2021.emnlp-main.241

    [13] Qi, F., Chen, Y., Zhang, X., Li, M., Liu, Z., & Sun, M. (2021). Mind the style of text! Adversarial and backdoor attacks based on text style transfer. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 4569–4580. https://doi.org/10.18653/v1/2021.emnlp-main.374

    [14] Gan, L., Li, J., Zhang, T., Li, X., Meng, Y., Wu, F., Yang, Y., Guo, S., & Fan, C. (2022). Triggerless backdoor attack for NLP tasks with clean labels. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2942–2952. https://doi.org/10.18653/v1/2022.naacl-main.214

    [15] Zeng, Y., Pan, M., Jahagirdar, H., Jin, M., Lyu, L., & Jia, R. (2023). Meta-Sift: How to sift out a clean subset in the presence of data poisoning? 32nd USENIX Security Symposium (USENIX Security 23), 1667–1684.

    https://www.usenix.org/conference/usenixsecurity23/presentation/zeng

    [16] Carlini, N., Tramèr, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, Ú., Oprea, A., & Raffel, C. (2021). Extracting training data from large language models. 30th USENIX Security Symposium (USENIX Security 21), 2633–2650. https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting

    [17] National Institute of Standards and Technology. (2023). Artificial intelligence risk management framework (AI RMF 1.0) (NIST AI 100-1). U.S. Department of Commerce. https://doi.org/10.6028/NIST.AI.100-1

    [18] Sandhu, R. S., Coyne, E. J., Feinstein, H. L., & Youman, C. E. (1996). Role-based access control models. Computer, 29(2), 38–47. https://doi.org/10.1109/2.485845

    [19] Servos, D., & Osborn, S. L. (2017). Current research and open problems in attribute-based access control. ACM Computing Surveys, 49(4), Article 65, 1–45. https://doi.org/10.1145/3007204

    [20] Rose, S., Borchert, O., Mitchell, S., & Connelly, S. (2020). Zero trust architecture (NIST Special Publication 800-207). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.800-207

    [21] Chandramouli, R., & Butcher, Z. (2023). A zero trust architecture model for access control in cloud-native applications in multi-cloud environments (NIST Special Publication 800-207A). National Institute of Standards and Technology.

    https://doi.org/10.6028/NIST.SP.800-207A

    [22] Wan, A., Wallace, E., Shen, S., & Klein, D. (2023). Poisoning language models during instruction tuning. Proceedings of the 40th International Conference on Machine Learning, 202, 35413–35425. https://proceedings.mlr.press/v202/wan23b.html

    [23] Lukas, N., Salem, A., Sim, R., Tople, S., Wutschitz, L., & Zanella-Béguelin, S. (2023). Analyzing leakage of personally identifiable information in language models. 2023 IEEE Symposium on Security and Privacy (SP), 346–363. https://doi.org/10.1109/SP46215.2023.10179300

    [24] Yang, Z., He, X., Li, Z., Backes, M., Humbert, M., Berrang, P., & Zhang, Y. (2023). Data poisoning attacks against multimodal encoders. Proceedings of the 40th International Conference on Machine Learning, 202, 39299–39313.

    https://proceedings.mlr.press/v202/yang23f.html

    [25] Zhu, B., Cui, G., Chen, Y., Qin, Y., Yuan, L., Fu, C., Deng, Y., Liu, Z., Sun, M., & Gu, M. (2023). Removing backdoors in pre-trained models by regularized continual pre-training. Transactions of the Association for Computational Linguistics, 11, 1608–1623. https://doi.org/10.1162/tacl_a_00622

    [26] Wang, W., & Feizi, S. (2023). Temporal robustness against data poisoning. Advances in Neural Information Processing Systems, 36. https://proceedings.neurips.cc/paper_files/paper/2023/hash/94bcb01789fccf15afe2764d8fe0f40e-Abstract-Conference.html

    [27] Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P. S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., Biles, C., Brown, S., Kenton, Z., Hawkins, W., Stepleton, T., Birhane, A., Hendricks, L. A., Rimell, L., Isaac, W., & Gabriel, I. (2022). Taxonomy of risks posed by language models. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 214–229. https://doi.org/10.1145/3531146.3533088

    [28] Kunaparaju, C. (2024). Adversarial AI in National Security: Understanding and Countering AI-Generated Cyber Threats. Letters in High Energy Physics.

    [29] Elnikety, E., Mehta, A., Vahldiek-Oberwagner, A., Garg, D., & Druschel, P. (2016). Thoth: Comprehensive policy compliance in data retrieval systems. 25th USENIX Security Symposium (USENIX Security 16), 637–654.

    https://www.usenix.org/conference/usenixsecurity16/technical-sessions/presentation/elnikety

    [30] Steinhardt, J., Koh, P. W., & Liang, P. S. (2017). Certified defenses for data poisoning attacks. Advances in Neural Information Processing Systems, 30, 3517–3529. https://proceedings.neurips.cc/paper/2017/hash/9d7311ba459f9e45ed746755a32dcd11-Abstract.html

    [31] Suciu, O., Marginean, R., Kaya, Y., Daumé III, H., & Dumitras, T. (2018). When does machine learning FAIL? Generalized transferability for evasion and poisoning attacks. 27th USENIX Security Symposium (USENIX Security 18), 1299–1316.

    https://www.usenix.org/conference/usenixsecurity18/presentation/suciu

    [32] Shokri, R., Stronati, M., Song, C., & Shmatikov, V. (2017). Membership inference attacks against machine learning models. 2017 IEEE Symposium on Security and Privacy (SP), 3–18. https://doi.org/10.1109/SP.2017.41

  • Downloads