Jurnal Sains dan Teknologi
Vol. 15 No. 1 (2026): April

Comparison of Text Classification Performance Using Manhattan Distance and Jaccard Distance in the KNN Algorithm

Ni Ketut Tari Tastrawati (Universitas Udayana, Denpasar, Indonesia)
Luh Sukma Mulyani (Universitas Udayana, Denpasar, Indonesia)
I Komang Gde Sukarsa (Universitas Udayana, Denpasar, Indonesia)
I Putu Eka Nila Kencana (Universitas Udayana, Denpasar, Indonesia)



Article Info

Publish Date
25 Apr 2026

Abstract

The rapid growth of short user-generated content has created significant challenges in accurately classifying sparse and high-dimensional text data, particularly in distinguishing spam from legitimate comments. Inappropriate selection of distance metrics and feature representations often leads to unstable and suboptimal classification performance. This study aims to analyze the effectiveness of Jaccard distance with Binary Bag of Words and Manhattan distance with TF-IDF in improving text classification performance. This research employs a quantitative approach with a comparative experimental design. The research subjects consist of 1,956 labeled text data points, including spam and non-spam comments. Data were collected using documentation techniques from a publicly available dataset, with the instrument in the form of labeled text data and a preprocessing pipeline to ensure data quality. Data analysis was conducted through K-Nearest Neighbors classification, performance evaluation using accuracy, precision, recall, and F1-score, and inferential statistical testing supported by effect size analysis. The results show that Jaccard distance consistently provides more stable and effective classification performance compared to Manhattan distance in handling sparse text data. This finding indicates that set-theoretic approaches are more suitable for high-dimensional text representations than geometric approaches. In conclusion, the alignment between distance metrics and feature representations plays a crucial role in optimizing classification performance. The implications of this study suggest that developers and researchers should prioritize appropriate metric-representation combinations to enhance the effectiveness and reliability of text classification systems.

Copyrights © 2026






Journal Info

Abbrev

JST

Publisher

Subject

Computer Science & IT Education

Description

Jurnal Sains dan Teknologi(JST) is a journal aims to be a peer-reviewed platform and an authoritative source of information. We publish original research papers, review articles and case studies focused on Mathematic, Biology, Physic, Chemistry, Informatic, Electronic and Machine as well as related ...