This Author published in this journals
All Journal Teknika
Rosihan
Department of Informatics, Faculty of Engineering, Universitas Khairun, Ternate, North Maluku, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Android-Based Research Title Similarity Detection Using a Combined Word2Vec and TF-IDF with Cosine Similarity Score Abyan Dzakwan Baksir; Rosihan; Muhammad Ridha Albaar
Teknika Vol. 15 No. 2 (2026): July 2026
Publisher : Center for Research and Community Service, Institut Informatika Indonesia (IKADO) Surabaya

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.34148/teknika.v15i2.1487

Abstract

Research title similarity detection is needed to support academic topic checking because manual title comparison is time-consuming and may fail to identify related topics expressed with different terms. This study develops an Android-based research title similarity detection system using a lightweight hybrid lexical-semantic approach that combines TF-IDF, Word2Vec Continuous Bag of Words (CBOW), and cosine similarity. The dataset was obtained through scraping from the GARUDA Portal and curated into 10,000 research titles related to computer science and information technology. The titles were processed through case folding, tokenization, stopword removal, and stemming using Sastrawi. TF-IDF was used to represent lexical term importance, while Word2Vec CBOW was used to capture contextual word relationships. The two similarity scores were integrated using weighted alpha configurations of 0.50, 0.60, and 0.70. The model was implemented in a Python FastAPI backend and tested through a Flutter-based Android application. Evaluation was conducted using 30 query titles with Top-5 retrieval results and manual relevance judgment based on a predefined 0–2 relevance rubric. The results show that TF-IDF only achieved the highest MAP@5 of 0.985972 and NDCG@5 of 0.939095. Among the hybrid configurations, alpha 0.70 produced the best performance with Precision@5 of 0.973333, MAP@5 of 0.981111, and NDCG@5 of 0.923408. These findings indicate that the hybrid model is competitive and more effective than Word2Vec only, while TF-IDF remains highly important for short research title matching.