Claim Missing Document
Check
Articles

Found 1 Documents
Search

Evaluation of Lexical and Semantic Representations in LexRank for Extractive Summarization of Indonesian Friday Sermon Texts Rusydiyyah, Taqiyyah Daaniys Shabrina; Supriyono, Supriyono; Aziz, Okta Qomaruddin; Ummah, Sofwatul; Rahmiasari, Siti Annisa
ILKOMNIKA Vol 8 No 2 (2026): Volume 8, Number 2, August 2026
Publisher : Lembaga Penelitian dan Pengabdian Masyarakat

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.28926/ilkomnika.v8i2.933

Abstract

Friday sermon transcripts from YouTube tend to be lengthy, repetitive, and inconsistently punctuated, creating a need for traceable extractive summarization that preserves source sentences. This study compares lexical TF-IDF representation and IndoBERT sentence embeddings in a continuous weighted LexRank architecture using 27 test documents separated at the document level from an initial corpus of 210 transcripts; 32 single-segment transcripts were excluded, leaving 178 eligible documents containing 16,131 sentences. TF-IDF LexRank, Semantic LexRank, and Lead-30% select 30% of the sentences from each document, and their outputs are compared with author-prepared and reviewed extractive reference summaries using ROUGE-1, ROUGE-2, ROUGE-L, 5,000 bootstrap iterations, paired Wilcoxon tests, and Holm correction. TF-IDF LexRank achieves the highest scores of 0.6633 on ROUGE-1, 0.5640 on ROUGE-2, and 0.5881 on ROUGE-L, significantly outperforming Semantic LexRank and Lead-30%. The findings indicate that, under this corpus and evaluation protocol, unigram-bigram lexical representation is better aligned with thematic term repetition and the extractive reference summaries than the tested semantic representation, while semantic representation still outperforms the positional baseline but does not achieve the best performance.