Luh Sukma Mulyani
Universitas Udayana, Denpasar, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Comparison of Text Classification Performance Using Manhattan Distance and Jaccard Distance in the KNN Algorithm Ni Ketut Tari Tastrawati; Luh Sukma Mulyani; I Komang Gde Sukarsa; I Putu Eka Nila Kencana
JST (Jurnal Sains dan Teknologi) Vol. 15 No. 1 (2026): April
Publisher : Universitas Pendidikan Ganesha

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.23887/jst-undiksha.v15i1.111518

Abstract

The rapid growth of short user-generated content has created significant challenges in accurately classifying sparse and high-dimensional text data, particularly in distinguishing spam from legitimate comments. Inappropriate selection of distance metrics and feature representations often leads to unstable and suboptimal classification performance. This study aims to analyze the effectiveness of Jaccard distance with Binary Bag of Words and Manhattan distance with TF-IDF in improving text classification performance. This research employs a quantitative approach with a comparative experimental design. The research subjects consist of 1,956 labeled text data points, including spam and non-spam comments. Data were collected using documentation techniques from a publicly available dataset, with the instrument in the form of labeled text data and a preprocessing pipeline to ensure data quality. Data analysis was conducted through K-Nearest Neighbors classification, performance evaluation using accuracy, precision, recall, and F1-score, and inferential statistical testing supported by effect size analysis. The results show that Jaccard distance consistently provides more stable and effective classification performance compared to Manhattan distance in handling sparse text data. This finding indicates that set-theoretic approaches are more suitable for high-dimensional text representations than geometric approaches. In conclusion, the alignment between distance metrics and feature representations plays a crucial role in optimizing classification performance. The implications of this study suggest that developers and researchers should prioritize appropriate metric-representation combinations to enhance the effectiveness and reliability of text classification systems.