Journal of Advanced Computer Knowledge and Algorithms
Vol. 3 No. 3 (2026): Journal of Advanced Computer Knowledge and Algorithms - July 2026

Enhancing Text Search Accuracy in Low-Resource Languages Using N-gram and Fuzzy Matching: A Case Study on Dari

Ahmad Ali Yaqin (Balkh University)
Sayed Ehsan Shamsi (Balkh University)
Khosrow Samadi (Balkh University)
Abdul Rahman Rahimi (Balkh University)



Article Info

Publish Date
20 Jun 2026

Abstract

Text search systems play an important role in modern information retrieval by providing fast and efficient access to digital information. However, despite the rapid development of digital technologies, these systems often encounter challenges when processing incomplete, misspelled, or orthographically inconsistent queries, particularly in low-resource languages such as Dari. The existence of spelling variations and limited linguistic resources reduces retrieval accuracy and negatively affects search performance. Information Retrieval (IR), as an important branch of Natural Language Processing (NLP), aims to retrieve accurate information from large text collections. Nevertheless, IR systems still face difficulties in low-resource languages due to limited datasets, insufficient computational resources, and the lack of language-specific retrieval tools. Therefore, improving search accuracy in the Dari language remains an important research challenge. This study aims to enhance text retrieval performance using the combination of n-gram analysis and Fuzzy Matching techniques. To achieve this objective, a mini search system was designed to process user queries and suggest the closest matching word whenever incomplete or misspelled inputs are detected. A Dari text dataset was collected and prepared through preprocessing, tokenization, and unique-word extraction. Fuzzy Matching was applied to measure similarity between the user query and dataset words, while n-gram analysis was used to examine structural similarity between words. The experimental findings demonstrated that the proposed mini search system successfully identified several incomplete and misspelled Dari words with relatively high similarity scores. These findings indicate that combining n-gram analysis with Fuzzy Matching can improve typo-tolerant text retrieval in low-resource languages such as Dari and provide a practical foundation for future research on Dari information retrieval systems.

Copyrights © 2026






Journal Info

Abbrev

jacka

Publisher

Subject

Computer Science & IT

Description

JACKA journal published by the Informatics Engineering Program, Faculty of Engineering, Universitas Malikussaleh to accommodate the scientific writings of the ideas or studies related to informatics science. JACKA journal published many related subjects on informatics science such as (but not ...