Mining Production Behavior Without Mining Customer Data: An AI Approach to Test Generation in Regulated Environments
Main Article Content
Abstract
Structural issues arise for regulated industries like banking, insurance, and healthcare where the need for accurate production data to verify the behaviour of software and the legal obligation to limit customer exposure to the software are in conflict. This paper combines recent developments in privacy-preserving data publishing, generative modelling, process mining, and search-based software testing into a single framework that allows one to mine production behaviour without mining customer data. It combines anonymisation primitives (like k-anonymity and l-diversity) with differential privacy techniques to limit the re-identification risk, uses generative adversarial architectures (conditioned tabular GANs) to produce statistically faithful synthetic event logs and injects the artefacts into whole-suite and mutation-guided test generation engines. Synthesized evidence shows that downstream analytical utility of differentially private generators is between 78 and 88 percent, while the residual re-identification risk is kept below 12 percent as compared to 31 to 42 percent for the classical k-anonymity and l-diversity approaches. Synthetic behavioral logs generate more than 90 percent branch coverage and 80 percent mutation score in 10 iterations of refinement. Federated learning also spreads the information of production patterns, but without transmitting raw data. Results indicate that privacy engineering alongside generative synthesis and search-based testing provides an appropriate and scalable alternative to production-data-dependent pipelines for quality assurance, which is audit and regulation compliant.
Article Details

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
References
[1] L. Sweeney, “k-Anonymity: A Model for Protecting Privacy,” Int. J. Uncertainty, Fuzziness Knowledge-Based Syst., vol. 10, no. 5, pp. 557–570, 2002, doi: 10.1142/S0218488502001648.
[2] K. El Emam and F. K. Dankar, “Protecting Privacy Using k-Anonymity,” J. Am. Med. Informatics Assoc., vol. 15, no. 5, pp. 627–637, 2008, doi: 10.1197/jamia.M2716.
[3] C. Dwork and A. Roth, “The Algorithmic Foundations of Differential Privacy,” Found. Trends Theor. Comput. Sci., vol. 9, no. 3–4, pp. 211–407, 2014, doi: 10.1561/0400000042.
[4] G. Fraser and A. Arcuri, “EvoSuite: Automatic Test Suite Generation for Object-Oriented Software,” in Proceedings of the 19th ACM SIGSOFT Symposium and the 13th European Conference on Foundations of Software Engineering, Association for Computing Machinery, 2011, pp. 416–419. doi: 10.1145/2025113.2025179.
[5] G. Fraser and A. Arcuri, “Whole Test Suite Generation,” IEEE Trans. Softw. Eng., vol. 39, no. 2, pp. 276–291, 2013, doi: 10.1109/TSE.2012.14.
[6] W. M. P. van der Aalst, Process Mining: Data Science in Action, 2nd ed. Springer Berlin, Heidelberg, 2016. doi: 10.1007/978-3-662-49851-4.
[7] S. A. Fahrenkrog-Petersen, H. van der Aa, and M. Weidlich, “PRIPEL: Privacy-Preserving Event Log Publishing Including Contextual Information,” in Business Process Management: BPM 2020, in Lecture Notes in Computer Science, vol. 12168. Springer, Cham, 2020, pp. 111–128. doi: 10.1007/978-3-030-58666-9_7.
[8] M. Kabierski, S. A. Fahrenkrog-Petersen, and M. Weidlich, “Privacy-Aware Process Performance Indicators: Framework and Release Mechanisms,” in Advanced Information Systems Engineering: CAiSE 2021, in Lecture Notes in Computer Science, vol. 12751. Springer, Cham, 2021, pp. 19–36. doi: 10.1007/978-3-030-79382-1_2.
[9] M. Rafiei and W. M. P. van der Aalst, “Privacy-Preserving Data Publishing in Process Mining,” in Business Process Management Forum: BPM Forum 2020, in Lecture Notes in Business Information Processing, vol. 392. Springer, Cham, 2020, pp. 122–138. doi: 10.1007/978-3-030-58638-6_8.
[10] I. J. Goodfellow et al., “Generative Adversarial Nets,” in Advances in Neural Information Processing Systems 27, Neural Information Processing Systems Foundation, 2014, pp. 2672–2680. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2014/hash/f033ed80deb0234979a61f95710dbe25-Abstract.html
[11] L. Xu, M. Skoularidou, A. Cuesta-Infante, and K. Veeramachaneni, “Modeling Tabular Data Using Conditional GAN,” in Advances in Neural Information Processing Systems 32, Neural Information Processing Systems Foundation, 2019, pp. 7333–7343. [Online]. Available: https://proceedings.neurips.cc/paper/2019/hash/254ed7d2de3b23ab10936522dd547b78-Abstract.html
[12] N. Papernot, M. Abadi, Ú. Erlingsson, I. J. Goodfellow, and K. Talwar, “Semi-Supervised Knowledge Transfer for Deep Learning from Private Training Data,” in Proceedings of the 5th International Conference on Learning Representations (ICLR 2017), OpenReview.net, 2017. [Online]. Available: https://openreview.net/forum?id=HkwoSDPgg
[13] N. C. Abay, Y. Zhou, M. Kantarcioglu, B. Thuraisingham, and L. Sweeney, “Privacy Preserving Synthetic Data Release Using Deep Learning,” in Machine Learning and Knowledge Discovery in Databases: ECML PKDD 2018, Proceedings, Part I, in Lecture Notes in Computer Science, vol. 11051. Springer, Cham, 2019, pp. 510–526. doi: 10.1007/978-3-030-10925-7_31.
[14] M. Abadi et al., “Deep Learning with Differential Privacy,” in Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, Association for Computing Machinery, 2016, pp. 308–318. doi: 10.1145/2976749.2978318.
[15] P. Kairouz et al., “Advances and open problems in federated learning,” Found. trends® Mach. Learn., vol. 14, no. 1–2, pp. 1–210, 2021, [Online]. Available: https://doi.org/10.1561/2200000083
[16] Q. Yang, Y. Liu, T. Chen, and Y. Tong, “Federated Machine Learning: Concept and Applications,” ACM Trans. Intell. Syst. Technol., vol. 10, no. 2, pp. 1–19, 2019, doi: 10.1145/3298981.
[17] A. Machanavajjhala, D. Kifer, J. Gehrke, and M. Venkitasubramaniam, “l-Diversity: Privacy Beyond k-Anonymity,” ACM Trans. Knowl. Discov. Data, vol. 1, no. 1, 2007, doi: 10.1145/1217299.1217302.
[18] E. Choi, S. Biswal, B. Malin, J. Duke, W. F. Stewart, and J. Sun, “Generating Multi-label Discrete Patient Records using Generative Adversarial Networks,” in Proceedings of the 2nd Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research, vol. 68. PMLR, 2017, pp. 286–305. [Online]. Available: https://proceedings.mlr.press/v68/choi17a.html
[19] C. Charitou, S. Dragicevic, and A. d’Avila Garcez, “Synthetic Data Generation for Fraud Detection using GANs,” 2021, arXiv. doi: 10.48550/arXiv.2109.12546.
[20] Y. Jia and M. Harman, “An Analysis and Survey of the Development of Mutation Testing,” IEEE Trans. Softw. Eng., vol. 37, no. 5, pp. 649–678, 2011, doi: 10.1109/TSE.2010.62.