Claim Missing Document
Check
Articles

Found 23 Documents
Search

Implementasi Metode TF-IDF dan Cosine Similarity pada Sistem Pencarian Artikel yang Relevan Selfira Madoa; Ida Mulyadi; Darniati Darniati
Journal of Muhammadiyah’s Application Technology Vol. 5 No. 2 (2026)
Publisher : Universitas Muhammadiyah Makassar

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.26618/a6wkdt68

Abstract

ABSTRAKPerkembangan teknologi informasi menyebabkan peningkatan volume data teks digital, khususnya artikel ilmiah, yang menuntut adanya sistem pencarian informasi yang mampu menyajikan hasil secara relevan dan kontekstual. Pencarian berbasis pencocokan kata kunci secara literal dinilai belum optimal dalam menangani variasi bahasa dan konteks kueri. Oleh karena itu, penelitian ini bertujuan untuk mengimplementasikan dan mengevaluasi metode Term Frequency–Inverse Document Frequency (TF-IDF) dan Cosine Similarity pada sistem pencarian artikel ilmiah berbahasa Indonesia. Penelitian ini menggunakan pendekatan kuantitatif dengan metode eksperimen, di mana data berupa judul dan abstrak artikel diperoleh dari repositori digital terbuka. Tahapan preprocessing teks meliputi case folding, tokenisasi, stopword removal, dan stemming untuk meningkatkan kualitas representasi data. Hasil penelitian menunjukkan bahwa sistem mampu menghasilkan nilai precision hingga 0,75 dan F1-score sebesar 0,67, yang mengindikasikan bahwa metode TF-IDF dan Cosine Similarity efektif dalam meningkatkan relevansi hasil pencarian. Dengan demikian, sistem yang dikembangkan mampu memberikan hasil pencarian yang lebih akurat dan kontekstual dibandingkan metode pencarian berbasis kata kunci literal, serta layak diterapkan pada repositori artikel ilmiah berskala kecil hingga menengah. Kata Kunci: TF-IDF, Cosine Similarity, Sistem Pencarian Informasi, Artikel Ilmiah, Text Mining ABSTRACTThe rapid growth of information technology has led to a significant increase in digital text data, particularly scientific articles, thereby requiring effective information retrieval systems capable of providing relevant and contextual results. Conventional keyword-based search methods are often insufficient in handling linguistic variations and complex query contexts. Therefore, this study aims to implement and evaluate the Term Frequency–Inverse Document Frequency (TF-IDF) and Cosine Similarity methods in an Indonesian scientific article search system. This research adopts a quantitative approach with an experimental method, using article titles and abstracts obtained from open-access digital repositories as the research dataset.Text preprocessing stages include case folding, tokenization, stopword removal, and stemming to improve data consistency and representation quality. The results indicate that the proposed system achieves a precision value of up to 0.75 and an F1-score of 0.67, demonstrating that the combination of TF-IDF and Cosine Similarity effectively enhances the relevance of search results. Thus, the developed system provides more accurate and contextual article retrieval compared to literal keyword matching and is suitable for implementation in small to medium-scale academic repositories. Keywords: TF-IDF, Cosine Similarity, Information Retrieval System, Scientific Articles, Text Mining
Penerapan Algoritma KNN dengan K-Fold Cross Validation Untuk Diagnosa Risiko Diabetes Mellitus Hafipa Sudiadarma; Ida Mulyadi; Fahrim Irhamna Rachman
Journal of Muhammadiyah’s Application Technology Vol. 5 No. 2 (2026)
Publisher : Universitas Muhammadiyah Makassar

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.26618/q2zkdm10

Abstract

ABSTRAK Diabetes mellitus (DM) merupakan penyakit kronis dengan prevalensi yang terus meningkat dan berpotensi menimbulkan berbagai komplikasi serius apabila tidak terdeteksi sejak dini. Keterbatasan metode diagnostik konvensional dalam menangani data kesehatan yang besar dan kompleks mendorong pemanfaatan pendekatan berbasis machine learning. Penelitian ini bertujuan untuk membangun model prediksi risiko diabetes mellitus menggunakan algoritma K-Nearest Neighbor (KNN) dengan metode Stratified K-Fold Cross Validation. Dataset yang digunakan terdiri dari 1.041 data pasien yang diperoleh dari Rumah Sakit Haji Makassar, dengan variabel meliputi usia, tekanan darah, status gula darah sewaktu, indeks massa tubuh, dan lingkar perut. Tahapan penelitian meliputi pemrosesan data, normalisasi menggunakan Standard Scaler, pemodelan KNN dengan metrik jarak Manhattan, serta evaluasi kinerja model menggunakan akurasi, precision, recall, dan F1-score. Hasil penelitian menunjukkan bahwa model KNN mampu mencapai rata-rata akurasi sebesar 87,55% dengan performa yang stabil pada setiap fold. Analisis feature importance menunjukkan bahwa tekanan darah sistolik, lingkar perut, dan gula darah sewaktu merupakan faktor yang paling berpengaruh terhadap status gula darah. Hasil ini menunjukkan bahwa algoritma KNN berpotensi digunakan sebagai alat bantu deteksi dini risiko diabetes mellitus berbasis data kesehatan. ABSTRACT Diabetes mellitus (DM) is a chronic disease with a continuously increasing prevalence and the potential to cause various serious complications if not detected early. The limitations of conventional diagnostic methods in handling large and complex health data have encouraged the use of machine learning-based approaches. This study aims to develop a diabetes mellitus risk prediction model using the K-Nearest Neighbor (KNN) algorithm with the Stratified K-Fold Cross Validation method. The dataset consisted of 1,041 patient records obtained from Haji Hospital Makassar, including variables such as age, blood pressure, random blood glucose level, body mass index, and waist circumference. The research stages included data preprocessing, normalization using Standard Scaler, KNN modeling with the Manhattan distance metric, and model performance evaluation using accuracy, precision, recall, and F1-score. The results showed that the KNN model achieved an average accuracy of 87.55% with stable performance across each fold. Feature importance analysis indicated that systolic blood pressure, waist circumference, and random blood glucose level were the most influential factors affecting blood glucose status. These findings suggest that the KNN algorithm has the potential to be used as a decision-support tool for the early detection of diabetes mellitus risk based on health data..
Klasifikasi Berita Hoaks pada Media Online Menggunakan Open Source Intelligence dan Algoritma Support Vector Machine Muh. Darmawan Aryadinata; Fachrim Irhamna Rachman; Ida Mulyadi
Journal of Muhammadiyah’s Application Technology Vol. 5 No. 2 (2026)
Publisher : Universitas Muhammadiyah Makassar

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.26618/gv4ndf79

Abstract

ABSTRAKPerkembangan media online yang pesat memudahkan masyarakat dalam memperoleh informasi secara cepat dan luas. Namun, kemudahan tersebut juga meningkatkan penyebaran berita hoaks yang dapat menimbulkan dampak negatif bagi masyarakat. Oleh karena itu, diperlukan suatu sistem yang mampu mengklasifikasikan berita hoaks secara otomatis dan akurat. Penelitian ini bertujuan untuk membangun sistem klasifikasi berita hoaks pada media online menggunakan metode Open Source Intelligence (OSINT) dan algoritma Support Vector Machine (SVM). Dataset diperoleh dari sumber terbuka berbasis web, yaitu situs klarifikasi hoaks dan portal berita online, yang kemudian melalui tahapan Preprocessing meliputi pembersihan teks, normalisasi, dan tokenisasi. Proses ekstraksi fitur dilakukan menggunakan metode Term Frequency–Inverse Document Frequency (TF-IDF) untuk merepresentasikan teks dalam bentuk numerik. Model klasifikasi dibangun menggunakan algoritma SVM dengan kernel linear karena efektif dalam menangani data teks berdimensi tinggi.Hasil pengujian menunjukkan bahwa model yang dikembangkan mampu mengklasifikasikan berita hoaks dan non-hoaks dengan tingkat akurasi sebesar 96,20%, serta didukung oleh nilai precision, recall, dan F1-score yang tinggi. Hal ini menunjukkan bahwa kombinasi metode OSINT, TF-IDF, dan SVM efektif dalam membangun sistem klasifikasi berita hoaks berbasis web dengan performa yang baik. Kata kunci: Klasifikasi Teks, Berita Hoaks, Media Online, OSINT, TF-IDF, Support Vector Machine. ABSTRACTThe rapid development of online media makes it easier for people to obtain information quickly and widely. However, this convenience also increases the spread of hoax news, which can have negative impacts on society. Therefore, a system capable of automatically and accurately classifying hoax news is needed. This research aims to develop a hoax news classification system in online media using Open Source Intelligence (OSINT) methods and the Support Vector Machine (SVM) algorithm. The Dataset was obtained from web-based open sources, namely hoax clarification websites and online news portals. The data were then subjected to Preprocessing stages including text cleaning, normalization, and tokenization. The feature extraction process was carried out using the Term Frequency–Inverse Document Frequency (TF-IDF) method to represent text in numerical form. The classification model was built using the SVM algorithm with a linear kernel because it is effective in handling high-dimensional text data. Test results show that the developed model is capable of classifying hoax and non-hoax news with an accuracy rate of 96.20%, supported by high precision, recall, and F1-score values. This demonstrates that the combination of OSINT, TF-IDF, and SVM methods is effective in building a web-based hoax news classification system with good performance. Keywords: Text Classification, Hoax News, Online Media, OSINT, TF-IDF, Support Vector Machine.