Abdul Rahman Rahimi
Balkh University

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Enhancing Text Search Accuracy in Low-Resource Languages Using N-gram and Fuzzy Matching: A Case Study on Dari Ahmad Ali Yaqin; Sayed Ehsan Shamsi; Khosrow Samadi; Abdul Rahman Rahimi
Journal of Advanced Computer Knowledge and Algorithms Vol. 3 No. 3 (2026): Journal of Advanced Computer Knowledge and Algorithms - July 2026
Publisher : Department of Informatics, Universitas Malikussaleh

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29103/jacka.v3i3.27644

Abstract

Text search systems play an important role in modern information retrieval by providing fast and efficient access to digital information. However, despite the rapid development of digital technologies, these systems often encounter challenges when processing incomplete, misspelled, or orthographically inconsistent queries, particularly in low-resource languages such as Dari. The existence of spelling variations and limited linguistic resources reduces retrieval accuracy and negatively affects search performance. Information Retrieval (IR), as an important branch of Natural Language Processing (NLP), aims to retrieve accurate information from large text collections. Nevertheless, IR systems still face difficulties in low-resource languages due to limited datasets, insufficient computational resources, and the lack of language-specific retrieval tools. Therefore, improving search accuracy in the Dari language remains an important research challenge. This study aims to enhance text retrieval performance using the combination of n-gram analysis and Fuzzy Matching techniques. To achieve this objective, a mini search system was designed to process user queries and suggest the closest matching word whenever incomplete or misspelled inputs are detected. A Dari text dataset was collected and prepared through preprocessing, tokenization, and unique-word extraction. Fuzzy Matching was applied to measure similarity between the user query and dataset words, while n-gram analysis was used to examine structural similarity between words. The experimental findings demonstrated that the proposed mini search system successfully identified several incomplete and misspelled Dari words with relatively high similarity scores. These findings indicate that combining n-gram analysis with Fuzzy Matching can improve typo-tolerant text retrieval in low-resource languages such as Dari and provide a practical foundation for future research on Dari information retrieval systems.