cover
Contact Name
Bambang
Contact Email
bambang.afriadi@yahoo.co.id
Phone
+6285692038195
Journal Mail Official
bambang.afriadi@yahoo.co.id
Editorial Address
Kompleks Universitas Negeri Jakarta, Kampus A, Gd. Muhammad Hatta Jl. Rawamangun Muka, Jakarta Timur, 13220, Indonesia
Location
Kota adm. jakarta timur,
Dki jakarta
INDONESIA
JURNAL EVALUASI PENDIDIKAN
ISSN : 20867425     EISSN : 26203073     DOI : https://doi.org/10.21009/JEP
Core Subject : Education,
JEP is the acronym of Jurnal Evaluasi Pendidikan. JEP is a national and international journal published two times in a year (March and October). It specializes in Instrument Development, Evaluation of Educational Programs, Assessment and Measurement in education. This journal is intended to communicate original research, policy brief and book review on the subject. JEP has been published since 2011 and has been using online system since 2017.
Articles 188 Documents
Computerized Adaptive Testing Berbasis IRT: Prinsip, Mekanisme, dan Simulasi Sederhana Maria Sumunaringtyas; Edi Istiyono; Lilin Rofiqotul Ilmi
Jurnal Evaluasi Pendidikan Vol. 17 No. 1 (2026): JURNAL EVALUASI PENDIDIKAN
Publisher : PROGRAM STUDI PENELITIAN DAN EVALUASI PENDIDIKAN

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/jep.v17i1.65078

Abstract

Transformasi asesmen dari bentuk linear menuju bentuk adaptif menjadi kebutuhan penting untuk meningkatkan ketepatan pengukuran dalam bidang pendidikan. Penelitian ini bertujuan untuk mengkaji mekanisme kerja Computerized Adaptive Testing (CAT) melalui pendekatan simulasi menggunakan Microsoft Excel. Metode yang digunakan adalah simulasi kuantitatif yang mengintegrasikan algoritma Maximum Likelihood Estimation (MLE) dan bank soal yang telah dikalibrasi menggunakan model Item Response Theory (IRT) 1-PL. Kebaruan penelitian ini terletak pada penerapan aturan penghentian tes (stopping rules) yang mengombinasikan kriteria presisi (SEM ≤ 0,30) dan kriteria stabilitas (∆SEM < 0,03) untuk mengoptimalkan durasi tes. Hasil simulasi menunjukkan bahwa mekanisme adaptif mampu mencapai konvergensi estimasi kemampuan peserta dengan tingkat akurasi yang tinggi hanya dalam rentang 12–15 butir soal. Temuan ini membuktikan efektivitas CAT dalam menghasilkan pengukuran yang efisien dan stabil, sekaligus memberikan gambaran yang transparan mengenai logika fungsional asesmen adaptif bagi para praktisi pendidikan.
UJI VALIDITAS DAN RELIABILITAS INSTRUMEN KOMPETENSI KERJA MAHASISWA Putri Anditasari; Yusi Riksa Y
Jurnal Evaluasi Pendidikan Vol. 17 No. 1 (2026): JURNAL EVALUASI PENDIDIKAN
Publisher : PROGRAM STUDI PENELITIAN DAN EVALUASI PENDIDIKAN

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/jep.v17i1.65431

Abstract

This study examines the validity and reliability of a student work competency instrument designed to measure underlying individual characteristics causally related to effective and superior performance. The instrument is developed based on the Spencer and Spencer competency model, which includes five dimensions: motive, trait, self-concept, knowledge, and skill. Measurement uses behavioral indicators distinguishing superior and threshold performance as indicators of students’ readiness to meet workplace demands. This study employed a descriptive quantitative approach involving 209 undergraduate students from higher education institutions as participants in an instrument try-out. Item validity was assessed using corrected item–total correlation, while reliability was examined using Cronbach’s Alpha. Results indicated that 48 of 60 tested items were valid. Reliability analysis produced a Cronbach’s Alpha coefficient of 0.945, indicating very high internal consistency. Overall, the findings demonstrate that the instrument is reliable and suitable for assessing student work competencies and supporting students’ readiness for career transition.
Dominasi Model CIPP dalam Evaluasi Program Pendidikan Indonesia: Studi Literatur Sistematis terhadap Model-model Evaluasi Sofia Edriati; Ambiyar Ambiyar; Fifi Yasmi; Liza Husnita
Jurnal Evaluasi Pendidikan Vol. 17 No. 1 (2026): JURNAL EVALUASI PENDIDIKAN
Publisher : PROGRAM STUDI PENELITIAN DAN EVALUASI PENDIDIKAN

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/jep.v17i1.65711

Abstract

Evaluasi program sangat krusial untuk menjamin kualitas pendidikan di Indonesia, namun sintesis komprehensif tentang model evaluasi yang paling efektif dan sering digunakan masih terbatas. Studi literatur sistematis ini bertujuan menganalisis penerapan model CIPP, Kirkpatrick, Logic Model, dan CIRO dalam evaluasi program pendidikan di Indonesia periode 2015-2025. Mengikuti pedoman PRISMA 2020, pencarian dilakukan pada database Scopus, menghasilkan 19 studi yang memenuhi kriteria inklusi. Temuan menunjukkan dominasi signifikan model CIPP sebagai kerangka kerja paling komprehensif untuk mengevaluasi program pendidikan yang kompleks, unggul dalam menangkap dimensi konteks, input, proses, dan produk program secara holistik. Model Kirkpatrick terbukti efektif untuk evaluasi pelatihan profesional dan capacity building dengan fokus pada perubahan perilaku individu. Tidak ditemukan penggunaan Logic Model dan CIRO melainkan model lain seperti Theory of Change, Bradley, Stake’s, CIPPO, dan DEM dalam literatur-literatur tersebut. Integrasi CIPP dengan teknik kuantitatif dan pendekatan hibrid dengan Kirkpatrick terbukti meningkatkan ketelitian dan kelengkapan evaluasi.
PENGARUH PEMANFAATAN ALAT EVALUASI ZEP QUIZ TERHADAP HASIL BELAJAR PAI SISWA KELAS XI SAINS Alya Rushafah; Santi Lisnawati; Salati  Asmahasanah
Jurnal Evaluasi Pendidikan Vol. 17 No. 1 (2026): JURNAL EVALUASI PENDIDIKAN
Publisher : PROGRAM STUDI PENELITIAN DAN EVALUASI PENDIDIKAN

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/jep.v17i1.68069

Abstract

This study aimed to determine the effect of using Zep Quiz as an evaluation tool on students’ learning outcomes in Islamic Religious Education (PAI). The research employed a quantitative approach with a quasi-experimental method using a nonequivalent control group design. The population consisted of all eleventh-grade science students at SMA Negeri 3 Kota Bogor, with a sample of 68 students divided into an experimental class and a control class. The research instruments included pretest and posttest multiple-choice tests and documentation. Data were analyzed using the Wilcoxon and Mann-Whitney tests with the assistance of SPSS 26. The findings showed that the experimental class experienced a significant increase in posttest scores after using Zep Quiz, while the control class showed no significant improvement. The statistical test results indicated a significance value of 0.000 < 0.05, meaning that Zep Quiz had a significant effect on students’ PAI learning outcomes.
Pengembangan Instrumen Tes Kemampuan Berpikir Kritis Pada Mata Pelajaran IPAS Sekolah Dasar Desi Dwi Prastiti; Endang Susilaningsih
Jurnal Evaluasi Pendidikan Vol. 17 No. 1 (2026): JURNAL EVALUASI PENDIDIKAN
Publisher : PROGRAM STUDI PENELITIAN DAN EVALUASI PENDIDIKAN

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/jep.v17i1.68437

Abstract

This study aimed to develop a FRISCO-oriented (Focus, Reason, Inference, Situation, Clarity, Overview) critical thinking test instrument for fourth-grade elementary students in IPAS learning on Indonesian cultural diversity material. The study employed a Research and Development (R&D) method using the ADDIE model. The population consisted of elementary school students, while the sample involved 20 fourth-grade students of SDN Kejawan Putih I/243 Surabaya selected through purposive sampling. Research instruments included validation sheets and a 20-item multiple-choice critical thinking test. Content validity was assessed by three subject experts and two classroom teachers. Findings showed a content validity score of 0.89 (very valid) and a reliability coefficient of 0.842 (high category). Empirical testing indicated that 80% of items had moderate difficulty and good discrimination power, demonstrating that the instrument is feasible, reliable, and effective for measuring students’ critical thinking skills.
Leveraging Local and Global EdTech Ecosystems to Strengthen Digital Literacy and Learning: A Systematic Review and Meta-Analysis Bambang Afriadi
Jurnal Evaluasi Pendidikan Vol. 17 No. 1 (2026): JURNAL EVALUASI PENDIDIKAN
Publisher : PROGRAM STUDI PENELITIAN DAN EVALUASI PENDIDIKAN

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/jep.v17i1.71016

Abstract

Background: Educational technology (EdTech) is commonly evaluated through adoption, access, or platform use, although these indicators do not establish digital-literacy development or improved learning. Evidence is also fragmented between locally contextualized platforms, globally scaled services, critical governance research, and intervention studies. Objective: This study synthesized how local or national and global EdTech ecosystems support digital literacy and learning, with particular attention to Indonesia, Malaysia, and Thailand, and quantitatively examined whether hardware- and software-focused interventions differed in their association with learning outcomes in low- and middle-income settings. Methods: A three-layer evidence synthesis was conducted in accordance with PRISMA 2020 principles. A supplied Scopus corpus of 78 records published from 2020 to 4 August 2026 was screened alongside 23 sources from the revised source manuscript and 15 intervention reports identified through citation chasing. Ninety-eight records were retained in the systematic evidence map, 38 sources informed the critical full-text or authoritative-source synthesis, and 30 hardware/software effect-size contrasts from 13 unique reports were included in a secondary random-effects meta-analysis. Effects were aggregated within publication using generalized least squares with an assumed within-publication correlation of 0.50, then pooled using restricted maximum likelihood and modified Hartung-Knapp intervals. Results: The qualitative evidence consistently supported a conditional hybrid-ecology model: local platforms contributed curricular, linguistic, and cultural alignment, while global platforms expanded content, authoring, interoperability, and collaboration. Digital-literacy gains depended on purposeful tasks, teacher competence, equitable access, accessibility, institutional support, and data governance. At publication level, hardware-focused interventions produced a near-zero pooled effect (d = -0.055, 95% CI -0.251 to 0.141; I² = 94.2%), whereas software-focused interventions showed a positive but imprecise effect (d = 0.217, 95% CI -0.024 to 0.458; I² = 90.7%). The modality difference was not definitive in meta-regression (p = 0.096). Contrast-level sensitivity analysis favored software (d = 0.245, 95% CI 0.092 to 0.397), but this estimate was less conservative because correlated outcomes were treated as independent. Conclusions: EdTech contributes through mediated practices rather than technological presence. The convergent evidence favors investment in pedagogically integrated software, teacher capacity, accessibility, and governance over hardware distribution alone. Because the quantitative evidence was highly heterogeneous, secondary, and indirect for digital literacy, the pooled estimates should be interpreted as directional rather than causal.  
Evaluating a Digital Literacy Program for Misinformation Discernment and Digital Civil-Rights Awareness Among Vocational-School Adolescents Fitri Fitri; Iqbal Syafrudin; Mitra Mustaricha
Jurnal Evaluasi Pendidikan Vol. 17 No. 1 (2026): JURNAL EVALUASI PENDIDIKAN
Publisher : PROGRAM STUDI PENELITIAN DAN EVALUASI PENDIDIKAN

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/jep.v17i1.71018

Abstract

Digital literacy programs are increasingly used to strengthen misinformation discernment and responsible participation in digital environments, yet their evaluation may be distorted when participants begin with near-perfect knowledge scores. This study evaluated the immediate outcomes of a digital literacy program delivered to vocational-school-aged adolescents in Nambo and Lulut, Bogor Regency, Indonesia. The program combined conceptual instruction, case-based identification of misinformation, guided use of verification tools, discussion of digital rights and responsibilities, and an invitation to participate as digital-literacy agents. A one-group pretest-posttest evaluation was conducted on 18 June 2026. Of 20 pretest and 20 posttest responses, 19 valid identity-matched pairs were analyzed. Participants were 15-18 years old (M = 16.42, SD = 0.69). The ten-item assessment measured knowledge of digital literacy, misinformation, source verification, digital civil rights, hate-speech responses, personal-data protection, and digital ethics. Mean scores increased from 98.42 (SD = 6.88) to 99.47 (SD = 2.29). The Wilcoxon signed-rank test did not indicate a statistically significant aggregate change (p = .317, r = .229), largely because 94.7% of participants had already achieved the maximum score at pretest. Eighteen scores remained at 100, whereas one participant improved from 70 to 90. Item-level gains occurred in appropriate responses to hate speech and understanding the purpose of digital literacy. The program is therefore best interpreted as maintaining high competence while producing targeted improvement, rather than generating a broad score increase. Future evaluations should use more demanding performance tasks, delayed posttests, and behavior-based indicators to improve measurement sensitivity and establish retention and transfer.
Analisis Butir Soal Asesmen Sumatif Akhir Tahun (ASAT) Mata Pelajaran Ekonomi Kelas X di SMAN Kalisat Eka Nurcahya; Sukidin Sukidin; Choirul Hudha
Jurnal Evaluasi Pendidikan Vol. 17 No. 1 (2026): JURNAL EVALUASI PENDIDIKAN
Publisher : PROGRAM STUDI PENELITIAN DAN EVALUASI PENDIDIKAN

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/jep.v17i1.66119

Abstract

The study aimed to analyze the quality of ASAT economics test items based on three aspects using a quantitative approach using Anates. A total of 30 multiple-choice questions with 356 student answers were analyzed in this study. Data were collected using document and interview methods. The results showed that 96.7% of the questions were valid, 56.7% were categorized as easy, and the effectiveness of distractors varied with a total of 46% effective distractors, and the overall questions were good based on the omit aspect. This study only provides recommendations for ideal test items based on these three aspects, while the final decision on follow-up on the ASAT questions is left to the economics teacher as the question compiler. The conclusion shows that most questions are suitable to be stored in the ASAT question bank with revisions to less effective distractors to improve the quality of the questions. Keywords item validity, level of difficulty, distractor effectiveness, economics summative assessment