Privacy-Preserving Data Mining Techniques for Sensitive Datasets
-
DOI:
https://doi.org/10.67228/30713498/IJADSMC-2021PI9R8TPublished 05-05-2021
Privacy-Preserving Data Mining, Sensitive Datasets, Differential Privacy, Secure Multi-Party Computation, Data Anonymization, Federated Learning Issue
Section
ArticlesHow to Cite
[1]S. Shah, “Privacy-Preserving Data Mining Techniques for Sensitive Datasets”, IJADSMC, vol. 4, no. 1, pp. 01–14, May 2021, doi: 10.67228/30713498/IJADSMC-2021PI9R8T.Abstract
The appraisal of Privacy-Preserving Data Mining (PPDM) has become a crucial research area in current times as a result of the incredible increase in the applications of data-driven applications that use sensitive data (health records, financial transactions, social networks, and governmental databases). Although the data mining techniques have been offering effective tools in the extraction of valuable knowledge, they facilitate great risks to personal privacy when they are applied to sensitive data. Unauthorized disclosure, inference attack, and breach of data has brought up serious ethical, legal, and regulatory issues. As a result, it is difficult to find the compromise between data utility and privacy protection. This essay outlines an extensive analysis of privacy ensuring data mining methods that allow secure privacy of sensitive data without compromising on the analysis accuracy. The paper systematically investigates the ways of anonymization, perturbation, cryptography, and hybrid privacy models. Besides that, newer privacy models include differential privacy, federated learning, and secure multi-party computation are discussed. A systematic approach is given to assess PPDM methods using privacy strength, data utility, computational complexity and scalability. The paper also reports on the findings of the experiments by comparing them, thus showing trade-offs between privacy and performance. The problems, problems under open research and direction are also discovered. The results highlight the lack of universal best practices because no single method is universally the best and the use of PPDM methods should be applied based on the application. The paper is intended to be a reference book of researchers and practitioners looking to have strong privacy preservation solutions such in sensitive data mining tasks.
References
[1] Samarati, P., & Sweeney, L. (1998). Protecting privacy when disclosing information: k-anonymity and its enforcement through generalization and suppression. Proceedings of the IEEE Symposium on Research in Security and Privacy.
[2] Sweeney, L. (2002). k-anonymity: A model for protecting privacy. International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems, 10(5), 557–570.
[3] Machanavajjhala, A., Gehrke, J., Kifer, D., & Venkitasubramaniam, M. (2007). l-diversity: Privacy beyond k-anonymity. ACM Transactions on Knowledge Discovery from Data, 1(1), 3.
[4] Li, N., Li, T., & Venkatasubramanian, S. (2007). t-closeness: Privacy beyond k-anonymity and l-diversity. Proceedings of the IEEE 23rd International Conference on Data Engineering.
[5] Fung, B. C. M., Wang, K., Chen, R., & Yu, P. S. (2010). Privacy-preserving data publishing: A survey of recent developments. ACM Computing Surveys, 42(4), 1–53.
[6] Agrawal, R., & Srikant, R. (2000). Privacy-preserving data mining. Proceedings of the ACM SIGMOD International Conference on Management of Data.
[7] Domingo-Ferrer, J., & Torra, V. (2005). Ordinal, continuous and heterogeneous k-anonymity through microaggregation. Data Mining and Knowledge Discovery, 11(2), 195–212.
[8] Verykios, V. S., Bertino, E., Fovino, I. N., Provenza, L. P., Saygin, Y., & Theodoridis, Y. (2004). State-of-the-art in privacy preserving data mining. ACM SIGMOD Record, 33(1), 50–57.
[9] Lindell, Y., & Pinkas, B. (2000). Privacy preserving data mining. Journal of Cryptology, 15(3), 177–206.
[10] Goldreich, O. (2004). Foundations of cryptography: Volume 2, basic applications. Cambridge University Press.
[11] Dwork, C. (2006). Differential privacy. Proceedings of the 33rd International Colloquium on Automata, Languages and Programming (ICALP).
[12] Dwork, C., McSherry, F., Nissim, K., & Smith, A. (2006). Calibrating noise to sensitivity in private data analysis. Proceedings of the Theory of Cryptography Conference.
[13] Dwork, C., & Roth, A. (2014). The algorithmic foundations of differential privacy. Foundations and Trends in Theoretical Computer Science, 9(3–4), 211–407.
[14] McMahan, B., Moore, E., Ramage, D., Hampson, S., & y Arcas, B. A. (2017). Communication-efficient learning of deep networks from decentralized data. Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS).
Downloads
How to Cite
[1]S. Shah, “Privacy-Preserving Data Mining Techniques for Sensitive Datasets”, IJADSMC, vol. 4, no. 1, pp. 01–14, May 2021, doi: 10.67228/30713498/IJADSMC-2021PI9R8T.