Claim Missing Document
Check
Articles

Boosting CNN Accuracy for Sundanese Script Recognition through Feature Extraction Techniques Pradana, Musthofa Galih; Khoirunnisa, Hilda
Journal of Applied Informatics and Computing Vol. 9 No. 6 (2025): December 2025
Publisher : Politeknik Negeri Batam

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30871/jaic.v9i6.11332

Abstract

Sundanese script is included in the cultural heritage in Indonesia, especially the culture in West Java. As a society that appreciates and preserves Indonesian culture and art, active participation can be realized through efforts to strengthen and preserve this script, one of which is by utilizing digital media. One of the technology-based digital media that can be used to preserve culture is image detection to make it easier to recognize Sundanese script. One of the models that can be used is the Convolutional Neural Network (CNN) with the MobileNetV2 architecture, with limited resources this architecture is able to produce good detection. This study applies the Convolutional Neural Network (CNN) algorithm with the MobileNetV2 architecture which will be tested with two main test scenarios, namely by applying feature extraction and without using feature extraction. The focus of this study will explore the influence and significance of the influence of feature extraction on the final results of image detection using the Convolutional Neural Network (CNN). The two feature extraction models used are Local Binary Pattern and Gray-Level Co-occurrence Matrix. These two feature extraction models will be tested with Sundanese script image data with data of 2,300 Sundanese script images. The results of this study show that the best results were obtained in the Convolutional Neural Network (CNN) with Gray-Level Co-occurrence Matrix (GLCM) with the best accuracy results at 93.8%. This is because the addition of the Gray-Level Co-occurrence Matrix (GLCM) is able to capture spatial texture statistics such as contrast, homogeneity, entropy, and correlation between pixel pairs. With these results, it can be concluded that in this study feature extraction has an effect and is able to increase the detection accuracy of the Convolutional Neural Network (CNN) model with the MobileNetV2 architecture in Sundanese script image data.
YOLOv8-Based Microplastic Detection and Quantification in River Water Microscopic Images Musthofa Galih Pradana; Retno Dwi Nyamiati; Husna Muizzati Shabrina; Muhammad Adrezo; Nurhuda Maulana
Journal of Applied Data Sciences Vol 7, No 2: May 2026
Publisher : Bright Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47738/jads.v7i2.1262

Abstract

Plastic particles with various size variations such as microplastics are environmental contaminants that are widely found in waters and have the potential to cause negative impacts. The process of identifying plastic particles using microscopic imagery manually takes a lot of time and considerable cost. In order to provide an alternative solution as part of early detection, microscopic image-based plastic particle detection was carried out with the YOLOv8 architecture, accompanied by an estimate of microplastic abundance in microplastic units per cubic meter. This study aims to develop and evaluate the detection of plastic particles in microscopic images of river water. This research dataset consists of 300 microscopic images taken from three river locations in Indonesia and annotated for model training and testing. The results of the evaluation showed that the proposed model had an aggregate performance value with a precision value of 0.786, recall of 0.66, and mAP@0.5 of 0.731. Additional test results show that with the addition of image resolution, the precision value can increase to 0.804 and the value mAP@0.5 increases to 0.762, even at the expense of computing time, which is also increasing. Extended scenario-based analysis showed that more than 87% of the detected objects fell into the category of small objects, affecting the localization sensitivity and variability of the estimated MPS value. This study also validated the results of object detection with FTIR-based laboratory tests using a full quantitative agreement between the model detection results and the identification of plastic particle materials at the sampling location level. The main contribution and findings of this study is an integrated evaluation framework for object detection, particle size characterization which is expected to be an alternative to the initial screening tool for plastic particle content.
Semantic clustering of scientific abstracts with transformer embeddings and traditional text representations Musthofa Galih Pradana; Pujo Hari Saputro; Ardhyansyah Mualo; Arbiati Faizah; Wahyuni Fithratul Zalmi
International Journal of Advances in Applied Sciences Vol 15, No 2: June 2026
Publisher : Institute of Advanced Engineering and Science

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.11591/ijaas.v15.i2.pp532-540

Abstract

The large and diverse quantity of scientific documents in the world, and specifically in Indonesia, makes the process of processing scientific document data an interesting study. One that represents the entire scientific document is through abstracts. The approach that can be done for the process of processing and grouping documents is to apply clustering. In this case, text-based clustering is currently heavily influenced by the feature representation of the text data used. Some popular representations of features are term frequency-inverse document frequency (TF-IDF) and count vectorizer, but they still have significant weaknesses in the context of understanding the meaning of natural language. To cover the drawbacks, it can use transformers or embedding types. In this study, several test scenarios will be carried out to obtain information and an overview of how to compare the effectiveness of the traditional TF-IDF model and the bidirectional encoder representations from transformers (BERT) and sentence bidirectional encoder representations from transformers (SBERT) embedding models in Indonesian-language scientific abstract clustering with several clustering models, such as k-means and agglomerative. The results of the study showed that the most effective clustering obtained was by using an embedding model of a combination of BERT and k-means, which was the most consistent with the most optimal number of clusters being 2 clusters.
Behavioral Analysis of Semantic Similarity Metrics under Transformer-Based Representations Musthofa Galih Pradana; Nindy Irzavika; Nurhuda Maulana; Syaila Ananta Karenina; Salma Ashiila Rabbani
Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi) Vol 10 No 2 (2026): April 2026
Publisher : Ikatan Ahli Informatika Indonesia (IAII)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29207/resti.v10i2.7490

Abstract

Knowledge extraction has several approaches such as traditional approaches that rely on lexical representation capabilities, one of which is TF-IDF whose implementation can be combined with several classic similarity metrics such as cosine similarity and dice coefficient similarity. In addition to applying the lexical representation approach, this study tries to apply it to a more modern type of representation, namely transformer-embedding-based contextual representation. The data used in this study is abstract document data of students' theses. The findings of the study show that contextual embedding changes the behavior of similarity values. The results of the analysis showed an average ranking shift of 6.70 positions. The test results showed a weak rating correlation value (Spearman = 0.22; Kendall = 0.146), and the high-ranking alignment measured with NDCG (0.97), which shows structural differences in the order of similarity between lexical and contextual representations. Other findings also show that the gap in ranking produced by the two representations used is quite far due to the difference in the mechanism and working pattern of the two representations that are far different, the selection of the type of representation must be on the characteristics of the data to be processed, if looking at the character of the text data in academic documents, the selection of contextual representations based on transformer embedding will be more suitable with contextual understanding to avoid the use of variations in words avoid plagiarism detection when applying a semantic-based representation approach.
KOMPARASI METODE NAÏVE BAYES DAN C4.5 DALAM KLASIFIKASI LOYALITAS PELANGGAN TERHADAP LAYANAN PERUSAHAAN Musthofa Galih Pradana; Pujo Hari Saputro
Indonesian Journal of Business Intelligence (IJUBI) Vol 3 No 1 (2020): Indonesian Journal of Business Intelligence (IJUBI)
Publisher : Universitas Alma Ata

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21927/ijubi.v3i1.1205

Abstract

Keberadaan pelanggan bagi jalannya sebuah usaha sangatlah penting. Pelanggan memiliki  kecenderungan yakni  untuk tetap lanjut berlangganan dengan perusahaan atau sebaliknya berhenti berlangganan. Salah satu teknik yang dapat digunakan untuk mengidentifikasi kecenderungan loyalitas pelanggan  adalah dengan klasifikasi data. Berdasarkan data pelanggan yang dimiliki perusahaan dapat dilakukan pengolahan data atau data mining dengan mengkelompokan pelanggan yang loyal dan yang tidak loyal. Ada banyak metode yang dapat diterapkan untuk klasifikasi data, diantaranya adalah algortima Naïve Bayes dan C4.5. Kedua metode ini menghasilkan akurasi yang berbeda ketika digunakan untuk proses klasifikasi data. Digunakan 2 skenario dalam proses pengujian kedua algoritma,  skenario membagi data dalam data testing dan training serta skenario pengujian menggunakan cross validation. Hasil kedua skenario ini menunjukan bahwa metode C4.5 lebih unggul dibandingkan dengan metode Naïve Bayes dengan akurasi skenario 1 sebesar 78,6086 % dan skenario 2 akurasi sebesar 78,61%. AbstrakKeberadaan pelanggan bagi jalannya sebuah usaha sangatlah penting. Pelanggan memiliki  kecenderungan yakni  untuk tetap lanjut berlangganan dengan perusahaan atau sebaliknya berhenti berlangganan. Salah satu teknik yang dapat digunakan untuk mengidentifikasi kecenderungan loyalitas pelanggan  adalah dengan klasifikasi data. Berdasarkan data pelanggan yang dimiliki perusahaan dapat dilakukan pengolahan data atau data mining dengan mengkelompokan pelanggan yang loyal dan yang tidak loyal. Ada banyak metode yang dapat diterapkan untuk klasifikasi data, diantaranya adalah algortima Naïve Bayes dan C4.5. Kedua metode ini menghasilkan akurasi yang berbeda ketika digunakan untuk proses klasifikasi data. Digunakan 2 skenario dalam proses pengujian kedua algoritma,  skenario membagi data dalam data testing dan training serta skenario pengujian menggunakan cross validation. Hasil kedua skenario ini menunjukan bahwa metode C4.5 lebih unggul dibandingkan dengan metode Naïve Bayes dengan akurasi skenario 1 sebesar 78,6086 % dan skenario 2 akurasi sebesar 78,61%.
KOMPARASI METODE SUPPORT VECTOR MACHINE DAN NAÏVE BAYES DALAM KLASIFIKASI PELUANG PENYAKIT SERANGAN JANTUNG Musthofa Galih Pradana; Pujo Hari Saputro; Dhina Puspasari Wijaya
Indonesian Journal of Business Intelligence (IJUBI) Vol 5 No 2 (2022): Indonesian Journal of Business Intelligence (IJUBI)
Publisher : Universitas Alma Ata

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21927/ijubi.v5i2.2659

Abstract

The death rate in the world per year is 17.9 million due to cardiovascular disease, including heart and blood vessel disorders. This needs to be given more attention to anticipate the possible risk of a heart attack. One of the contributions in the field of technology to provide useful information about the risk of heart disease is by using a data processing approach or data mining technique by classifying the vulnerability to heart disease risk. The classification method used is Support Vector Machine and Naïve Bayes. The classification method will be carried out in a comparative process and the method that has the best accuracy will be sought. The scenarios used are 2 test scenarios, namely dividing the training data by 20% in scenario 1 and 40% in scenario 2. The final results of the research obtained are the best accuracy in the Support Vector Machine with scenario 1 of 87%.
Analisis Performa Algoritma Convolutional Neural Networks Menggunakan Arsitektur LeNet dan VGG16 Musthofa Galih Pradana; Hilda Khoirunnisa
Indonesian Journal of Business Intelligence (IJUBI) Vol 6 No 2 (2023): Indonesian Journal of Business Intelligence (IJUBI)
Publisher : Universitas Alma Ata

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21927/ijubi.v6i2.3765

Abstract

Identifying a person's self-identity can be done by recognizing facial images, where faces can often represent a person's identity. Facial identification with technology can benefit the effectiveness efficiency and accuracy of data. This identification process can be used with the help of algorithms that will check digital images with the necessary detection results. One algorithm that can be applied in classifying and detecting gender through facial image algorithms is Convolutional Neural Networks. Convolutional Neural Network algorithms have various architectures that have advantages in each architecture. This study compared the process of identifying a person's face to obtain information in the form of gender. The models compared in this study are the LeNet model and the VGG16 model. The identification and detection process was carried out using 800 photos for data training with gender labeling data and 240 photos for testing data. A comparison of these two models is necessary to get the best final model result. The final results obtained from this study the best accuracy of both architectures was obtained in the VGG16 architecture which reached an average accuracy of 100 in several epochs compared to the VGG16 architecture at 0.925 in the 46th epoch. This is due to a Rectified Linear Unit (ReLU) on the VGG16 architecture which can minimize errors and saturation.
Deteksi Kemiripan Dokumen Menggunakan Cosine Similarity Berdasarkan Representasi Teks Count Vectorizer Dan TF IDF Musthofa Galih Pradana; Nindy Irzavika; Nurhuda Maulana
Indonesian Journal of Business Intelligence (IJUBI) Vol 7 No 2 (2024): Indonesian Journal of Business Intelligence (IJUBI)
Publisher : Universitas Alma Ata

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21927/ijubi.v7i2.5170

Abstract

Tujuan mata kuliah skripsi atau tugas akhir menumbuhkan budaya berpikir kritis, dan menunjukan kemampuan untuk memecahkan permasalahan dengan konstruksi logis dari penelitian. Akan tetapi, dari banyaknya manfaat tersebut, ada beberapa permasalahan yang juga muncul dikarenakan mata kuliah ini. Plagiarisme adalah masalah umum. Mengambil karya orang lain, termasuk pendapat mereka sendiri, dan membuatnya seperti karya sendiri adalah plagiarisme. Langkah pertama dalam penggunaan teknologi adalah mendeteksi kesamaan dokumen sejak dini. Dalam hal ini, dokumen yang harus dikumpulkan oleh mahasiswa selama proses pengajuan judul skripsi mereka adalah abstrak. Ketika digunakan, algoritma cosine similarity adalah algoritma yang efisien secara komputasi karena sangat mudah dipahami dan dapat digunakan dengan data berskala besar. Penelitian ini dilakukan dengan dua pendekatan representasi teks yaitu dengan menggunakan TF-IDF dan Count Vectorizer. Data korpus yang digunakan dalam penelitian ini adalah 1600 data dokumen abstrak skripsi mahasiswa, dengan pengujian menggunakan 30 data untuk melihat kinerja algoritma cosine similarity dalam mendeteksi kesamaan dokumen abstrak. Hasil penelitian menunjukkan bahwa pendekatan representasi teks TF-IDF mendapatkan kesamaan di angka 7,72861 dan Count Vectorizer mendapatkan hasil di angka 16,85541 atau punya gap sebesar 9,1268 dengan keunggulan Count Vectorizer. Hal ini disebabkan Count Vectorizer menghitung frekuensi kata tanpa mempertimbangkan apakah kata tersebut umum atau jarang, sehingga kata-kata umum tetap berkontribusi penuh terhadap similarity.
Penguatan UMKM Ramah Lingkungan melalui Sinergi Inovasi Produk dan Transformasi Digital Musthofa Galih Pradana; I Wayan Rangga Pinastawa; Retno Dwi Nyamiati; Fadhli Suko Wiryanto; Ryan Setya Budi; Nurul Afifah Arifuddin; Muhammad Adrezo; Zatin Niqotaini
Jurnal Pengabdian kepada Masyarakat Bidang Ilmu Komputer Vol 4 No 1 (2025): Jurnal Pengabdian Kepada Masyarakat Bidang Ilmu Komputer (ABDIKOM)
Publisher : Universitas Pembangunan Nasional "Veteran" Jakarta

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.52958/abdikom.v4i1.12855

Abstract

Kawasan Industri Sentolo merupakan lokasi yang strategis bagi UMKM Jangkang Indah Craft, produsen kerajinan berbahan serat agel yang merupakan sumber daya alam khas Kabupaten Kulon Progo. Memasuki tahun 2025, UMKM ini dihadapkan pada dua tantangan utama, yaitu kesiapan melakukan pemasaran internasional secara mandiri serta penerapan teknik pewarnaan ramah lingkungan guna mendukung keberlanjutan ekosistem dan meminimalkan dampak ekologis. Menjawab kebutuhan tersebut, kegiatan Pengabdian kepada Masyarakat dirancang untuk memperkuat kapasitas pengelolaan usaha secara holistik. Program ini menerapkan metode Participatory Action Research meliputi sosialisasi, pendampingan, dan praktik langsung pada beberapa aspek kunci, yakni peningkatan manajemen pemasaran melalui pelatihan digital marketing berbasis website serta pemanfaatan platform e-Bay. Di sisi lain, peserta juga menerima pelatihan teknis mengenai teknik pewarnaan berkelanjutan, seperti pemanfaatan spirulina untuk menghasilkan warna hijau serta penggunaan ekstrak warna kuning dengan kandungan kimia yang lebih rendah. Pendampingan dilakukan dalam tiga sesi, dan hasil evaluasi melalui kuesioner menunjukkan tingkat kepuasan rata-rata 4,27 dari skala 1 - 5, yang mencerminkan kategori penilaian baik. Temuan ini sejalan dengan perkembangan positif dari pendampingan tahap sebelumnya, di mana keberadaan website terbukti meningkatkan branding awal UMKM dan membantu menjalin sejumlah kemitraan melalui kemudahan akses informasi.
Construction of a Dialect-Sensitive Javanese Semantic Lexicon to Support Machine Translation Systems Musthofa Galih Pradana; Ridwan Raafi'udin; Nurul Afifah Arifuddin; Mohammad Asaduzzaman Rasel
Journal of Applied Informatics and Computing Vol. 10 No. 4 (2026): August 2026
Publisher : Politeknik Negeri Batam

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30871/jaic.v10i4.13379

Abstract

The development of linguistic resources for natural language processing (NLP) in Javanese remains limited, especially regarding the representation of semantic relationships between different speech levels. This study aims to construct a Javanese semantic lexicon that integrates Indonesian lexical equivalents with three Javanese speech levels: ngoko, krama alus, and krama inggil. A research design based on lexical resource construction was employed, using a Javanese digital dictionary as the primary data source. The methodology included data extraction, preprocessing, semantic lexicon construction, analysis of speech level variation, and a preliminary exploration of polysemous lexical entries using automatic identification, followed by validation by native speakers. The resulting semantic lexicon successfully represents lexical relationships between levels in a structured manner. Analysis of speech-level variation revealed that partially distinct lexical patterns were the most dominant, with 733 entries, followed by fully distinct patterns (193 entries) and identical patterns (21 entries). These findings indicate that speech-level differences in Javanese are selectively realized and should be explicitly considered in the development of linguistic resources. Furthermore, preliminary exploration of polysemous candidates demonstrated that dictionary-based automatic identification can overestimate polysemy without linguistic validation. Only a limited number of lexical entries exhibited features consistent with genuine polysemous relationships. This study provides an initial basis for the development of Javanese semantic resources that are sensitive to speech-level variation and semantic complexity. The constructed semantic lexicon has the potential to support future research in NLP applications in Javanese, including politeness identification, lexical normalization, word sense disambiguation, and machine translation.