Andi Nurkholis
Universitas Pembangunan Nasional Veteran Yogyakarta

Published : 2 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 2 Documents
Search

Comparative Analysis of Email Spam Detection Using SVM with TF-IDF and Word2Vec on Multilingual Datasets Kaifa Ahlal Katamsyi; Ahmad Taufiq Akbar; Andi Nurkholis; Hari Prapcoyo; Bagus Muhammad Akbar; Shoffan Saifullah
Paradigma - Jurnal Komputer dan Informatika Vol. 28 No. 1 (2026): March 2026 Period
Publisher : LPPM Universitas Bina Sarana Informatika

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.31294/p.v28i1.12339

Abstract

The rapid growth of email communication has increased the prevalence of spam emails, which can disrupt productivity and compromise information security. This study presents a comparative analysis of two text representation methods—TF-IDF and Word2Vec—for spam email classification using a Support Vector Machine (SVM) with a Radial Basis Function kernel. The experiments utilized Indonesian and English email datasets totaling 5,421 emails, split into 75% training and 25% testing sets. Two scenarios were evaluated: baseline with default parameters and after hyperparameter optimization using Grid Search combined with K-Fold Cross Validation. The results indicate that TF-IDF consistently outperformed Word2Vec across both languages, achieving the highest accuracy of 0.9562 on the English dataset after tuning. Word2Vec showed substantial improvement following parameter adjustment, reducing the performance gap with TF-IDF. The findings highlight the importance of hyperparameter optimization for enhancing the quality of feature representations and improving classification performance. This study also demonstrates that TF-IDF provides more stable results across different linguistic contexts, while Word2Vec benefits significantly from careful tuning. The results provide practical insights for implementing efficient spam email detection systems in multilingual environments. Future research could explore additional classifiers, deep learning approaches, and contextual embeddings to further improve classification accuracy and robustness.
MODEL KLASIFIKASI PENERIMA PROGRAM KARTU INDONESIA PINTAR MENGGUNAKAN METODE XGBOOST Andi Nurkholis; Styawati Styawati; Susi Susi; Alifah Chairul Munawar
INTI Nusa Mandiri Vol. 20 No. 2 (2026): INTI Periode Februari 2026
Publisher : Lembaga Penelitian dan Pengabdian Pada Masyarakat

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.33480/inti.v20i2.7770

Abstract

Poverty is one of the major factors contributing to the low quality of education in Indonesia. The Smart Indonesia Program (Program Indonesia Pintar) is a cash assistance program distributed through the Smart Indonesia Card (Kartu Indonesia Pintar/KIP) to support students from economically disadvantaged families, ranging from elementary to higher education levels. This study aims to classify students who are eligible to receive the Smart Indonesia Card using the XGBoost method, with a case study conducted at SMPN 02 Kebun Tebu. The dataset used in this study consists of independent and dependent variables. The independent variables include father’s age, mother’s age, father’s education level, mother’s education level, father’s income, mother’s income, number of family dependents, and students’ academic average scores. The dependent variable is the eligibility status of KIP recipients as the target class. Two classification models were developed using data split ratios of 70:30 and 80:20. The model with a 70:30 data split achieved an accuracy of 0.9048, a precision of 0.9034, a recall of 0.9072, and an F1-score of 0.9053. Meanwhile, the model with an 80:20 data split demonstrated better performance, with an accuracy of 0.9167, a precision of 0.9149, a recall of 0.9189, and an F1-score of 0.9169. The optimal model obtained from this study can be utilized by schools to support policy decision-making in determining eligible Smart Indonesia Card recipients, ensuring that educational assistance is distributed accurately, adequately, and equitably