Fieter Brain Pasaribu
Universitas Pendidikan Ganesha

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Synthetic Data Quantity in Animal Classification Fieter Brain Pasaribu; Luh Joni Erawati Dewi; I Nyoman Saputra Wahyu Wijaya; Ni Wayan Marti; Ni Putu Novita Puspa Dewi
Jurnal Teknologi Informasi dan Pendidikan Vol. 20 No. 1 (2027): Jurnal Teknologi Informasi dan Pendidikan
Publisher : Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/jtip.v20i1.1137

Abstract

The need for large datasets remains a challenge in deep learning-based classification models such as Convolutional Neural Networks. Limited data can lead to overfitting, where the model performs well on training data but poorly on testing data. To address this, numerous studies have been conducted using both traditional augmented data and synthetic data generated by Generative Adversarial Networks (GANs). GANs have proven effective across a wide range of dataset domains. However, no research has focused on the effect of the amount of synthetic data on model classification accuracy, leading to confusion about how much should be used to avoid overfitting. Therefore, this study will review the effectiveness of GANs by examining the effect of synthetic data ranging from small amounts to 20 times the original dataset size. The synthetic image results from GANs were then tested using them as supplementary datasets for MobileNetV2 and EfficientNetB0. The results show that the effectiveness of synthetic data fluctuates significantly. At certain amounts, the use of synthetic data can improve accuracy, such as in MobileNetV2, which increased from 73% to 77%. However, at certain amounts, the use of synthetic data also decreases accuracy from 73% as the baseline to 66%. This fluctuating condition is caused by the synthetic data from GANs not being consistent in representing the original data well. The lack of quality of GAN is caused by the architecture used, which is still too simple to capture data patterns with a wide variety of perspectives