Structural issues arise for regulated industries like banking, insurance, and healthcare where the need for accurate production data to verify the behaviour of software and the legal obligation to limit customer exposure to the software are in conflict. This paper combines recent developments in privacy-preserving data publishing, generative modelling, process mining, and search-based software testing into a single framework that allows one to mine production behaviour without mining customer data. It combines anonymisation primitives (like k-anonymity and l-diversity) with differential privacy techniques to limit the re-identification risk, uses generative adversarial architectures (conditioned tabular GANs) to produce statistically faithful synthetic event logs and injects the artefacts into whole-suite and mutation-guided test generation engines. Synthesized evidence shows that downstream analytical utility of differentially private generators is between 78 and 88 percent, while the residual re-identification risk is kept below 12 percent as compared to 31 to 42 percent for the classical k-anonymity and l-diversity approaches. Synthetic behavioral logs generate more than 90 percent branch coverage and 80 percent mutation score in 10 iterations of refinement. Federated learning also spreads the information of production patterns, but without transmitting raw data. Results indicate that privacy engineering alongside generative synthesis and search-based testing provides an appropriate and scalable alternative to production-data-dependent pipelines for quality assurance, which is audit and regulation compliant.
Copyrights © 2024