cover
Contact Name
Hendra Kurniawan
Contact Email
hendra.kurniawan@darmajaya.ac.id
Phone
-
Journal Mail Official
jodmapps@darmajaya.ac.id
Editorial Address
Jl. Z.A. Pagar Alam No. 93 Gedong Meneng, Bandar Lampung Lampung
Location
Kota bandar lampung,
Lampung
INDONESIA
Journal of Data Science Methods and Applications
ISSN : -     EISSN : 30905605     DOI : https://doi.org/10.30873/jodmapps
Theoretical Foundations: Architecture, Management and Process for Data Science Artificial Intelligence Classification and Clustering Data Pre-Processing, Sampling and Reduction Deep Learning Educational Data Mining Forecasting High Performance Computing for Data Analytics Learning Classifiers Learning Theory Optimization Methods Probabilistic and Statistical Models and Theories Scientific Data and Big Data Analytics Statistical Learning Machine Learning and Knowledge Discovery: Big Data Visualization, Modeling and Analytics Data and Knowledge Visualization Database Technology Knowledge Based Neural Networks Knowledge Discovery (Heterogeneous, Unstructured and Multimedia Data) Knowledge Discovery in Network and Link Data Knowledge Discovery in Social Networks Learning for Streaming Data Machine Learning for High-Performance Computing Multimedia/Stream/Text/Visual Analytics Spatial/Temporal Data Computational Data Science: Big Data Computational for Big Data Analysis Computational Intelligence for Pattern Recognition and Medical Imaging Computer Application for Data Analytics Computer Architecture for Data Analytics Computer Graphics for Data Analytics Data Acquisition, Integration, Cleaning Data Visualizations Data Wrangling Databases Decision Making from İnsights, Hidden Patterns Intelligent Information Retrieval Optimization for Data Analytics Probabilistic And İnformation-Theoretic Methods Search and Mining Support Vector Machines Time Series Analysis Applications: Bioinformatics Applications Biomedical Informatics Applications Biometrics Applications Collaborative Filtering Applications Data and Information Semantics Applications Data Mining Algorithms Applications Data Mining Systems Applications Data Streams Mining Applications Database and Information System Performance Applications Database Systems & Applications Electronic Commerce and Web Technologies Applications Electronic Government & E-participation Applications Graph Mining Applications Healthcare Applications Image Analysis Applications Information Retrieval Applications Multimedia Data Mining Applications Natural Language Processing Applications Pre-Processing Techniques Applications Spatial Data Mining Applications Statistical and Scientific Databases Applications Web Search Applications
Articles 21 Documents
Segmentasi Pelanggan Berdasarkan Kebutuhan Primer Skunder dan Tersier Menggunakan K-Means Clustering Indah, Caesaliana Indah Mu’assyaroh; Arkan, M. Rizieq Sultan; Galuh, Galuh Sitoresmi; Sabrina, Sabrina Rizkiya; Zida, Zida Nadhifah Aulia Kencana
Journal of Data Science Methods and Applications Vol. 1 No. 2 (2025)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

In the digital marketing era, companies are required to deeply understand customer behavior in order to develop targeted strategies. Customer segmentation is a common technique used to group customers based on similarities in their characteristics and consumption behaviors. This study aims to identify customer segments using unsupervised learning techniques with the K-Means clustering algorithm. The dataset, obtained from Kaggle, contains 2,240 customer records with demographic and purchase behavior attributes. The six primary features analyzed include Income, Age, TotalChildren, MntMeatProducts, NumCatalogPurchases, and Recency. The clustering results reveal distinct customer groups with different characteristics and purchasing tendencies, which can be used to develop more personalized and efficient marketing strategies.
Prediksi Tingkat Stres Mahasiswa Selama Pembelajaran Daring Menggunakan Algoritma Machine Learning Rocky, Rocky Khalifah Akbar; dede; Aldo, Aldo Septian Raharjo; Celvin, Celvin Immanuel Suhendar
Journal of Data Science Methods and Applications Vol. 1 No. 2 (2025)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

The sudden transition to online learning during the COVID-19 pandemic has had a significant psychological impacton students, particularly in the form of increased stress levels. This study aims toidentify and analyze the factors that influence student stress during online learningusing a quantitative approach and predictive modeling. Data were obtained from 100 students aged 18–25 years, covering variables such as screen time, sleep duration, physical activity, pre-exam anxiety, and changes inacademic performance. Statistical analysis showed that high screen time, less than 6 hours of sleep, andacademic anxiety were significantly associated with increased stress levels (p < 0.01). The Random Forest modelsuccessfully predicted stress categories with 82% accuracy and identified sleep duration as the mostdominant factor. These findings indicate the need for more adaptive academic policy reforms regarding mental health,including digital load management, healthy sleep education, and the integration of psychological support. This studyprovides an empirical basis for educational institutions to design data-driven preventive interventions toreduce the prevalence of stress among students.
Prediksi Jenis Ancaman Siber Global Menggunakan Algoritma Random Forest Pratama, Yudista Nanda; Ali, Rionaldi; Aris, Noki; Saputra, Agus; Saputra, M . David
Journal of Data Science Methods and Applications Vol. 1 No. 2 (2025)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Global cyber threats continue to increase along with the widespread digital transformation across various sectors. This study aims to predict the types of cyberattacks based on global historical data from 2015 to 2024. Data was obtained from the Global Cybersecurity Threats dataset, which includes information on countries, affected sectors, and types of attacks. The method used was supervised learning with the Random Forest algorithm, which is known to be effective for classifying and analyzing complex variables. The results show that this algorithm is capable of identifying attack patterns with high accuracy and assisting in early threat detection. This research is expected to contribute to the development of data-driven cybersecurity systems and predictive modeling. Global cyberthreats continue to grow in complexity, along with society's increasing reliance on digital systems. This study aims to analyze trends and predict the types of cyberthreats based on historical data from 2015 to 2024, obtained from the Global Cybersecurity Threats dataset. The method used was supervised learning with the Random Forest algorithm to classify attack types and predict potential financial losses. The analysis results show that the model can identify important patterns that can assist organizations in mitigating cyber risks. This research contributes to the development of data-driven cyber threat intelligence systems
Prediksi Pengunduran Diri Karyawan Menggunakan Metode Algoritma Random Forest Prasetyo, Bima Restu; Apiliani, Lusy Pebi; Intan, Citra Nur; Jonathan, Kenny
Journal of Data Science Methods and Applications Vol. 1 No. 2 (2025)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Employee attrition is a critical issue in human resource management as it directly affects a company’s productivity and operational efficiency. Therefore, a data-driven prediction system is needed to identify potential employee resignation risks at an early stage. This study aims to build an employee attrition classification model using the Random Forest algorithm, implemented in the RapidMiner software. The dataset used in this study is derived from the IBM HR Analytics Employee Attrition Dataset. The research process includes data cleaning, attribute transformation, model building, and performance evaluation using a confusion matrix and metrics such as accuracy, precision, and recall. The results show that the Random Forest model achieved an accuracy of 91.04%, a precision of 100% for the “Yes” class, and a recall of 44.37%. Furthermore, it was found that the variables JobLevel and TotalWorkingYears significantly influence attrition status. Therefore, this model can serve as a decision support tool in identifying employee attrition risks and designing more effective, data-driven retention strategies
Evaluasi dan Perbandingan Metode XGBoost dan LightGBM Dalam Deteksi Dini Penyakit Alzheimer Muhammad Rezky Adytama; Egi Safitri; Asmaul Dwi Akbar; Nicholas Svensons; Raka Sebastian Musin
Journal of Data Science Methods and Applications Vol. 2 No. 1 (2026)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Early detection of Alzheimer’s disease is a key step in slowing disease progression and improving patients’ quality of life. This study evaluates and compares the performance of the XGBoost and LightGBM algorithms in diagnosing Alzheimer’s disease, using a longitudinal dataset comprising 2.149 subjects. The dataset underwent meticulous preprocessing, including handling missing data, feature selection, and duplicate data removal, to ensure data reliability. Model evaluation was conducted based on accuracy, precision, recall, F1-score, and Mean Squared Error (MSE) metrics. The results demonstrate that the LightGBM algorithm outperforms XGBoost, achieving an accuracy of 87%, precision of 87%, recall of 85%, F1-score of 85%, and an MSE of 0.12. The advantages of LightGBM include computational efficiency and the ability to handle large-scale data, making it more effective than XGBoost. This study provides guidance on selecting the optimal algorithm for clinical applications, enabling more effective early interventions to mitigate the adverse impacts of Alzheimer’s disease.
Analisis Klasifikasi Multikelas Obesitas Menggunakan Algoritma Decision Tree Classifier, Random Forest Classifier, dan Support Vector Classifier (SVC) Septa, Oon; Triyasri, Novita; Permata, Maharani Aulia; Salsabila, Aghitsna; Firdani, Fahri
Journal of Data Science Methods and Applications Vol. 2 No. 1 (2026)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Obesity remains a significant global health challenge, making early classification and detection essential to minimize the risk of more serious degenerative diseases. This study compares three machine learning algorithms—Random Forest, Decision Tree, and Support Vector Classifier (SVC) to determine which model is most effective in predicting weight status categories and obesity risks. The dataset used includes physical features such as age, height, weight, and Body Mass Index (BMI). The research process encompasses data preprocessing stages, feature correlation analysis, data splitting, model training, and evaluation using various performance metrics.The research results indicate that Random Forest demonstrates the highest discriminative ability with an AUC of 0.99, showing perfect accuracy in distinguishing between obesity categories. Decision Tree provides identical results in terms of accuracy at 95.45%, but with a slightly lower AUC value of 0.96. Meanwhile, SVC yields competitive results with an accuracy of 81.82%, although its performance remains below the two tree-based models. Overall, this study demonstrates that ensemble methods such as Random Forest hold great potential for use as decision support systems in detecting and classifying obesity status more accurately and reliably.
Prediksi Tingkat Pengetahuan Mahasiswa Menggunakan Logistic Regression dan Random Forest Triyasri, Novita; Andini, Rekha Apriliana; Chandra, Aurea Ivana; Saprianti, Assyifa; Yusiandra, Erick
Journal of Data Science Methods and Applications Vol. 2 No. 1 (2026)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

particularly for predicting students’ knowledge levels more accurately. This study aims to compare the performance of Logistic Regression and Random Forest algorithms in predicting students’ knowledge levels using the User Knowledge Modeling dataset. The dataset consists of 258 instances with five numerical attributes, namely STG, SCG, STR, LPR, and PEG, and one target variable UNS representing students’ knowledge levels. The reserch stages include data selection, preprocessing, data normalization, train-test splitting, and handling class imbalance using the SMOTE method. Model performance is evaluated using accuracy, precision, recall, and F1-score metrics. The results show that Logistic Regression outperforms Random Forest, achieving higher accuracy and F1-score values. These findings indicate that the relationships among variables in the dataset tend to be linear. Therefore, Logistic Regression is considered more suitable for predicting students’ knowledge levels in this study.
Prediksi Survivabilitas Pasien Kanker Payudara dengan Penanganan Imbalance Data Menggunakan Algoritma Machine Learning Fikri, Ruki Rizal Nul; Prasetyo, Indra; Soleh, Ary Sofyan; Pratama, Reza Lintang Hana; Kurniawan, Hendra
Journal of Data Science Methods and Applications Vol. 2 No. 1 (2026)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Breast cancer is one of the leading causes of death among women worldwide. A major challenge in modeling patient survivability prediction is imbalanced data, where the number of surviving patients significantly outweighs the deceased ones. This study aims to compare the performance of three machine learning algorithms: Logistic Regression, Support Vector Classifier (SVC), and Gradient Boosting Classifier, in predicting patient survivability status. To address the class imbalance issue, Random Over Sampling (ROS) technique was applied during the data preprocessing stage. The methodology includes categorical data encoding, resampling, and model evaluation using accuracy, precision, recall, and F1-score metrics. Experimental results show that the application of ROS successfully balanced the class distribution. Among the three models tested, the Gradient Boosting algorithm demonstrated the best performance compared to linear and vector-based models. This study provides insights into the importance of handling imbalanced data to improve the accuracy of AI-based medical diagnoses.
DIAGNOSIS PCOS BERDASARKAN FAKTOR GAYA HIDUP DAN FAKTOR REPRODUKSI MENGGUNAKAN REGRESI LOGISTIK DAN RANDOM FOREST Kurniawan, Hendra; Kultsum, Rahil Urwa; Safitri, Egi; Antonio, Yandi Jaya; Andini, Rekha Aprilia; Syahputra, Lingga; Adytama, Muhammad Rezky
Journal of Data Science Methods and Applications Vol. 2 No. 1 (2026)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Polycystic Ovary Syndrome (PCOS) is a common endocrine disorder occurring in women of reproductive age, with a global prevalence ranging from 6% to 21%. Current management of PCOS remains limited to symptomatic treatment without addressing the root cause. This study aims to build an accurate predictive model for PCOS diagnosis in Indonesia by analyzing lifestyle and reproductive factors using machine learning algorithms, such as Logistic Regression and Random Forest.The research dataset consists of 541 patient records, which were divided into 80% for training and 20% for testing. The data was normalized using the Min-Max Scaler method, and class imbalance was handled using the SMOTE (Synthetic Minority Oversampling Technique) method. The models were validated using the K-Fold Cross-Validation method and evaluated based on accuracy, precision, recall, and F1-score.The results showed that Logistic Regression with SMOTE in predicting reproductive factors achieved the highest accuracy (82%), while Random Forest with SMOTE demonstrated more stable performance based on average accuracy, particularly for reproductive factors. ROC curve analysis also revealed that Logistic Regression with SMOTE in predicting reproductive factors achieved the highest AUC ($0.84$), making the Logistic Regression model superior in predicting the diagnosis compared to Random Forest. This study confirms that reproductive factors play a more dominant role in predicting PCOS compared to lifestyle factors. Utilizing machine learning algorithms can effectively predict PCOS to support management and prevention, as well as accelerate the early detection process of PCOS.
Analisis Feature Importance pada Penyakit Alzheimer Menggunakan Random Forest Puteri Yuni, Sundari; Oktavianingrum, Margareta; Aksa, Fadhilla
Journal of Data Science Methods and Applications Vol. 2 No. 1 (2026)
Publisher : Program Studi Sains Data - Institut Informatika dan Bisnis Darmajaya

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Alzheimer’s disease is a progressive neurodegenerative disorder characterized by cognitive decline and impaired daily functioning, particularly among the elderly population. The increasing global prevalence of Alzheimer’s disease highlights the need for accurate, efficient, and accessible early detection methods. This study aims to analyze feature importance in predicting Alzheimer’s disease using the Random Forest algorithm. The dataset used is secondary data obtained from Kaggle, consisting of 2,148 patient records with 35 features covering demographic, medical, cognitive, and functional aspects. The research methodology includes data preprocessing, class imbalance handling using SMOTE, feature selection with SelectKBest, and model training and evaluation using Random Forest with K-Fold cross-validation. The results demonstrate that the Random Forest model achieved excellent performance with an accuracy of 94%, and balanced precision and recall values of 0.94. Feature importance analysis reveals that Functional Assessment, Activities of Daily Living (ADL), and Mini-Mental State Examination (MMSE) are the most influential predictors of Alzheimer’s disease. These findings indicate that cognitive and functional indicators play a more significant role in early Alzheimer’s detection than other medical factors. This study is expected to contribute to the development of effective and interpretable medical decision support systems for early Alzheimer’s disease detection.

Page 2 of 3 | Total Record : 21