The need for large datasets remains a challenge in deep learning-based classification models such as Convolutional Neural Networks. Limited data can lead to overfitting, where the model performs well on training data but poorly on testing data. To address this, numerous studies have been conducted using both traditional augmented data and synthetic data generated by Generative Adversarial Networks (GANs). GANs have proven effective across a wide range of dataset domains. However, no research has focused on the effect of the amount of synthetic data on model classification accuracy, leading to confusion about how much should be used to avoid overfitting. Therefore, this study will review the effectiveness of GANs by examining the effect of synthetic data ranging from small amounts to 20 times the original dataset size. The synthetic image results from GANs were then tested using them as supplementary datasets for MobileNetV2 and EfficientNetB0. The results show that the effectiveness of synthetic data fluctuates significantly. At certain amounts, the use of synthetic data can improve accuracy, such as in MobileNetV2, which increased from 73% to 77%. However, at certain amounts, the use of synthetic data also decreases accuracy from 73% as the baseline to 66%. This fluctuating condition is caused by the synthetic data from GANs not being consistent in representing the original data well. The lack of quality of GAN is caused by the architecture used, which is still too simple to capture data patterns with a wide variety of perspectives
Copyrights © 2027