Claim Missing Document
Check
Articles

Klasifikasi Smartphone Berdasarkan Spesifikasi Menggunakan Decision Tree C4.5 dan Random Forest dengan Hyperparameter Tuning Akbar Ainurrofik; Ilham Saifudin; Rosita Yanuarti
Indonesian Journal of Multidisciplinary on Social and Technology Vol. 4 No. 3 (2026): Juli - Oktober
Publisher : PT Ilmu Data Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.69693/ijmst.v4i3.13271

Abstract

The rapid growth of smartphone products has produced substantial variation in technical specifications and price segments, making manual grouping increasingly complex. This study identifies the most informative specification using the C4.5 Decision Tree and develops a smartphone price-segment classification model using Random Forest optimized through hyperparameter tuning. The dataset contains 3,260 smartphone records and six input features: RAM capacity, internal storage, number of processor cores, battery capacity, main-camera resolution, and screen size. Prices in Indian rupees were converted into Indonesian rupiah and used only to generate three class labels: Entry-Level, Mid-Range, and Flagship; price was excluded from model inputs to prevent data leakage. The dataset was divided using a stratified 80:20 split. C4.5 analysis identified internal storage as the root node at a threshold of 192 GB with a Gain Ratio of 0.3804. The baseline Random Forest achieved 85.12% accuracy, 82.92% macro precision, 80.81% macro recall, and 81.77% macro F1-score. GridSearchCV with five-fold StratifiedKFold selected 100 trees, max_depth 20, min_samples_split 2, min_samples_leaf 1, max_features sqrt, and bootstrap=True. The tuned model maintained 85.12% accuracy while improving macro precision to 83.44%, macro recall to 81.03%, and macro F1-score to 82.15%. The tuned model was selected because it produced a better balance across classes and was subsequently implemented in a web-based classification system.
Sistem Rekomendasi Produk Fitness dan Olahraga Berbasis Content-Based Filtering Menggunakan Doc2vec dan Cosine Similarity Airlangga Zabaniyah Putro Irdianto; Ilham Saifudin; Taufiq Timur W.
Indonesian Journal of Multidisciplinary on Social and Technology Vol. 4 No. 3 (2026): Juli - Oktober
Publisher : PT Ilmu Data Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.69693/ijmst.v4i3.13272

Abstract

Recommendation systems help users identify products that match their needs amid the growing number of choices on digital platforms. The increasing variety of fitness and sports products can cause information overload and make relevant products difficult to find. This study develops and evaluates a content-based filtering recommendation system using Doc2Vec and cosine similarity. The study applies a quantitative experimental method to the Amazon Reviews'23 Dataset. Product_name, category, and clean_text are used as content attributes, while Product_id is used as the product identity. The data undergo selection, text cleaning, case folding, tokenization, and stopword removal, followed by an 80% training and 20% testing split. The Doc2Vec model uses the Paragraph Vector Distributed Bag-of-Words architecture with 100-dimensional vectors. Query vectors are produced through infer_vector and compared with product vectors using cosine similarity to create Top-N recommendations. Evaluation on four queries shows average Precision of 0.9500, 0.9000, and 0.8750; average Recall of 0.1100, 0.2078, and 0.6082; and average NDCG of 0.9732, 0.9584, and 0.9361 for Top-5, Top-10, and Top-30, respectively. The results show that smaller Top-N values provide higher precision and ranking quality, while larger Top-N values retrieve a broader set of relevant products. The system is implemented as a Flask-based website named Bolang.