This study developed and evaluated an Indonesian thesis-title similarity detection system to support the early screening of potentially similar research titles. The reference corpus consisted of 28,117 unique thesis titles collected from a publicly accessible institutional repository. The system integrates TF-IDF-based lexical retrieval, SBERT-based semantic retrieval, Union Top-K candidate selection, and machine learning reranking. Logistic Regression, XGBoost, and LightGBM were trained using threshold-derived pseudo-labels, while the final model performance was evaluated on an independent human-validated dataset constructed from a title-disjoint evaluation partition. A total of 900 candidate title pairs were independently assessed by two academic validators, achieving a Cohen’s kappa of 0.877. LightGBM achieved the best performance, with an accuracy of 94.67% and a macro F1-score of 0.9467. ISO/IEC 25010 evaluation showed a 100% functional pass rate, an 84.17% usability score, performance scores of 92.3% on mobile and 96.8% on desktop, while the automated security assessment categorized the staging deployment as Fairly Secure with two medium-risk findings. The system can support academic title-similarity screening; however, its generalizability remains limited by the use of a single-institution corpus and requires further multi-institutional validation.
Copyrights © 2026