Synthetic Data Generation Using Generative Artificial Intelligence

Authors

  • Adedokun Taofeek Ladoke Akintola University of Technology image/svg+xml Author

Keywords:

Synthetic Data Generation, Generative AI, Privacy Preservation, Data Augmentation, Model Fidelity

Abstract

The increasing demand for large-scale, high-quality datasets for machine learning has been constrained by privacy regulations, data scarcity, and high annotation costs. Generative Artificial Intelligence (GenAI) has emerged as a transformative solution through synthetic data generation—artificially created data that mimics real-world statistical properties while safeguarding sensitive information. This study investigates the effectiveness of various synthetic data generation methods across three critical dimensions: fidelity, utility, and privacy. Using a mixed-methods design combining quantitative evaluation with application-specific assessment, we examined three primary generation approaches—statistical-based methods (synthpop), deep learning methods (Generative Adversarial Networks, Variational Autoencoders), and Large Language Model-based generation—across multiple datasets and downstream tasks. Quantitative results demonstrate that classical methods such as synthpop and copula consistently achieve strong performance across both high-dimensional and low-dimensional datasets, combining low re-identification risk (ε-identifiability of 0.25-0.35) with competitive utility. Deep learning models achieve higher distributional fidelity but incur elevated privacy risks (ε-identifiability > 0.4). Notably, general utility metrics showed only weak correlation with specific task performance, emphasizing the need for multimetric evaluation. The findings reveal that data normalization techniques critically impact generation quality, with Max-Abs normalization consistently yielding more accurate and stable synthetic profiles across all models. These findings contribute to a nuanced understanding of synthetic data generation and provide practical guidelines for selecting appropriate methods based on application requirements, data characteristics, and privacy constraints.

References

1.

Akpabli, G., Yeboah, N. N., Gyimah, E., Agyei, E., Kojo, A. B., Rahnema, H., & Okyere Williams, M. (2026, June). Machine learning enhanced geomechanical characterization and wellbore stability assessment in the offshore South Tano Field, Ghana. Paper presented at the ARMA US Rock Mechanics/Geomechanics Symposium (p. D022S045R001). ARMA.

2.

Dankwa-Mullan, I. (2024). Health equity and ethical considerations in using artificial intelligence in public health and medicine. Preventing Chronic Disease, 21, Article 240245. https://doi.org/10.5888/pcd21.240245

3.

de Vries-Gao, A. (2025). The carbon and water footprints of data centers and what this could mean for artificial intelligence. Patterns, 7(1), Article 101430. https://doi.org/10.1016/j.patter.2025.101430

4.

Han, Y., Wu, Z., Li, P., Wierman, A., & Ren, S. (2024). Health-informed computing: Estimating and addressing the public health impact of data centers (arXiv:2412.06288). arXiv. https://doi.org/10.48550/arXiv.2412.06288

5.

Lee, J. T., Liu, V. T. N., Ali, S., Chen, P.-C., Lee, C.-C., Huang, C.-W., Li, V. C.-S., Chen, H.-H., Duh, W.-J., Marthias, T., & Atun, R. (2025). The impact of artificial intelligence on the health economy, workforce productivity, and administrative efficiency: A systematic review. medRxiv. https://doi.org/10.1101/2025.10.05.25337345

6.

Nguyen, T. T. (2026a). AI-powered precision public health: A national framework for targeting chronic

7.

disease prevention and early intervention. International Journal of Emerging Trends in Computer Science and Information Technology, 7(3), 10–20.

8.

Nguyen, T. T. (2026b). Transforming prior authorization through artificial intelligence: A national framework for reducing administrative burden and improving patient access. International Journal of AI, BigData, Computational and Management Studies, 7(3), 17–27.

9.

Yeboah, N. N., Agyei, E., Gyimah, E., Ayensigna, A. A., Okyere Williams, M., Akpabli, G.,... & Vidzro, B. (2026, June). Petrophysical characterization and reservoir quality evaluation of the Three Forks geothermal reservoir, Williston Basin. Paper presented at the ARMA US Rock Mechanics/Geomechanics Symposium (p. D032S048R012). ARMA.

10.

Seetala, S. R. (2025). Architecting autonomous data platforms: Integrating AI-driven governance, metadata intelligence, and data mesh principles. International Journal of Science, Engineering and Technology, 13(1).

11.

Vasa, M. R. (2024). Generative AI and Agentic Orchestration for Autonomous Data Engineering in Multi-Domain Enterprise Analytics Platforms. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 5(4), 364-369.

12.

Mudusu, S. K. (2025). Data Engineering Challenges in AI-Driven Healthcare IT Systems: Navigating Real-Time Analytics and Interoperability.Kesarpu, S. (2025). Zero-Trust Architecture in Java Microservices. International Journal of Networks and Security, 5(01), 202-214.

13.

Narapareddy, V. S. R., & Yerramilli, S. K. (2022). Scaling the ServiceNow CMDB for distributed infrastructures. International Journal of Engineering Technology Research & Management, 6(10), 101–113.

14.

Gollapudi, R. (2022). Risk-controlled near-zero-downtime Oracle database migration using GoldenGate. International Journal of Computational and Experimental Science and Engineering, 8(3), 113–123. https://doi.org/10.22399/ijcesen.5382

Downloads

Published

2026-09-09

Similar Articles

21-30 of 58

You may also start an advanced similarity search for this article.