Bulletin of Computer Science Research
Vol. 6 No. 4 (2026): June 2026

Analisis Efektivitas IndoBERT untuk Klasifikasi Multilabel Terjemahan Hadis Bukhari Menggunakan Logistic Regression

Achmad Yamin Harahap (Universitas Islam Negeri Sultan Syarif Kasim Riau, Pekanbaru)
Nazruddin Safaat H (Universitas Islam Negeri Sultan Syarif Kasim Riau, Pekanbaru)
Surya Agustian (Universitas Islam Negeri Sultan Syarif Kasim Riau, Pekanbaru)
Suwanto Sanjaya (Universitas Islam Negeri Sultan Syarif Kasim Riau, Pekanbaru)
Teddie D (Universitas Islam Negeri Sultan Syarif Kasim Riau, Pekanbaru)



Article Info

Publish Date
30 Jun 2026

Abstract

Hadith serves as the second source of guidance after the Quran, directing Muslims in various aspects of life; the *Sahih al-Bukhari* collection is among the most renowned. The complex nature of their meanings often encompassing multiple categories of messages poses a significant challenge for manual text classification, particularly as data volume grows. In this study, the content of the hadith often includes multiple message types, such as recommendations, prohibitions, and general information. This research aims to evaluate an automated classification system for Indonesian translations of *Sahih al-Bukhari* hadith, categorizing them into three classes: Information, Recommendation, and Prohibition. The study is motivated by the vast number of hadith, which requires significant time and deep understanding for people to grasp the core message of each one. This classification system is intended to facilitate the identification of primary messages, thereby making the processes of searching, studying, and understanding hadith more effective and efficient. IndoBERT is employed to generate contextual vector representations capable of capturing deeper semantic meaning, while Logistic Regression is selected for its efficiency and stability with high-dimensional data. Evaluation is conducted using a train-validation-test split approach, alongside accuracy and macro F1-score metrics. The study achieved an average F1-score of 67.43%, demonstrating that the combination of IndoBERT and Logistic Regression yields strong, consistent classification performance for this multi-label task.

Copyrights © 2026






Journal Info

Abbrev

bulletincsr

Publisher

Subject

Computer Science & IT

Description

Bulletin of Computer Science Research covers the whole spectrum of Computer Science, which includes, but is not limited to : • Artificial Immune Systems, Ant Colonies, and Swarm Intelligence • Bayesian Networks and Probabilistic Reasoning • Biologically Inspired Intelligence • Brain-Computer ...