Enrico Budi Santoso
Universitas Jenderal Achmad Yani, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Identification of Hoax News in the Using Community TF-RF and C5.0 Tree Decision Algorithm Enrico Budi Santoso; Yulison Herry Chrisnanto; Gunawan Abdillah
Enrichment: Journal of Multidisciplinary Research and Development Vol. 1 No. 6 (2023): Enrichment: Journal of Multidisciplinary Research and Development
Publisher : International Journal Labs

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.55324/enrichment.v1i6.58

Abstract

News has a great influence on social and political conditions, and the rapid circulation of information through social media increases the risk that people receive and redistribute hoax news. Identifying hoax news is therefore important to support the circulation of reliable information, particularly political news. This research aims to create a system for identifying hoax news using TF-RF feature weighting and the C5.0 Decision Tree algorithm and to evaluate its classification performance. The study uses 1,000 news data obtained by web scraping with the keywords "election 2024", "politics", and "checkfaktapilkadamafindo" from Turnbackhoax.id and Detik.com. The processing stages include preprocessing, TF-RF word weighting, division of training and test data, C5.0 classification, and evaluation using a confusion matrix. Three training/test scenarios were evaluated. The 70/30 scenario produced 79.33% accuracy, 80.50% precision, and 97.01% recall; the 80/20 scenario produced 79.50% accuracy, 81.32% precision, and 95.48% recall; and the 90/10 scenario produced 72.00% accuracy, 74.39% precision, and 89.71% recall. Among the tested scenarios, the 80/20 split provided the highest accuracy. These findings show that the combination of TF-RF weighting and C5.0 can be implemented as an automatic classification approach for political hoax-news identification, while performance remains dependent on the composition of the training and testing data