Nathaniel Clarence Haryanto
Department of Informatics, Universitas Kristen Duta Wacana

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Temu kembali dokumen sumber rujukan dalam sistem daur ulang teks Nathaniel Clarence Haryanto; Lucia Dwi Krisnawati; Antonius Rachmat Chrismanto
Jurnal Teknologi dan Sistem Komputer Volume 8, Issue 2, Year 2020 (April 2020)
Publisher : Department of Computer Engineering, Engineering Faculty, Universitas Diponegoro

Show Abstract | Download Original | Original Source | Check in Google Scholar | Full PDF (613.945 KB) | DOI: 10.14710/jtsiskom.8.2.2020.140-149

Abstract

The architecture of the text-reuse detection system consists of three main modules, i.e., source retrieval, text analysis, and knowledge-based postprocessing. Each module plays an important role in the accuracy rate of the detection outputs. Therefore, this research focuses on developing the source retrieval system in cases where the source documents have been obfuscated in different levels. Two steps of term weighting were applied to get such documents. The first was the local-word weighting, which has been applied to the test or reused documents to select query per text segments. The tf-idf term weighting was applied for indexing all documents in the corpus and as the basis for computing cosine similarity between the queries per segment and the documents in the corpus. A two-step filtering technique was applied to get the source document candidates. Using artificial cases of text reuse testing, the system achieves the same rates of precision and recall that are 0.967, while the recall rate for the simulated cases of reused text is 0.66.