I Gede Indra Aryasa
Universitas Diponegoro, Semarang, Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Assessment Practices in Physics Learning: A Systematic Literature Review of Functions, Instruments, and the Emerging Role of Technology (2015–2025) I Made Astra; I Gede Indra Aryasa
Jurnal Penelitian & Pengembangan Pendidikan Fisika Vol. 12 No. 1 (2026): JPPPF (Jurnal Penelitian dan Pengembangan Pendidikan Fisika), Volume 12 Issue
Publisher : Program Studi Pendidikan Fisika Universitas Negeri Jakarta, LPPM Universitas Negeri Jakarta, HFI Jakarta, HFI

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.21009/1.12109

Abstract

Assessment is the pivot on which physics teaching turns, yet the field has expanded so rapidly along several largely independent fronts , misconception diagnosis, formative feedback, higher-order thinking measurement, and, more recently, artificial-intelligence-supported scoring , that a consolidated picture has been difficult to obtain. This study reports a systematic literature review, conducted in accordance with the PRISMA 2020 guideline, of how assessment is conceived and operationalised in physics learning. Searches of Scopus, Web of Science, and ERIC returned 1,284 records, of which 38 studies published between 2015 and 2025 met the eligibility criteria after screening and quality appraisal. Extracted data were synthesised narratively around four review questions concerning assessment functions, instrument formats and validation strategies, the constructs being measured, and the penetration of digital and AI-based tools. The evidence suggests that diagnostic and formative purposes dominate the literature, that multi-tier diagnostic tests and rubric-scored open responses are the most frequently developed instruments, and that classical test theory remains the prevailing validation paradigm although Rasch and item-response approaches are gaining ground. Conceptual understanding and the diagnosis of misconceptions, together with higher-order and critical thinking, are the constructs most often targeted; affective and self-regulatory outcomes remain comparatively neglected. A small but fast-growing cluster of work applies large language models to automated grading and feedback, with reported human–machine agreement that is encouraging yet uneven across question types. The review argues that the field would benefit from tighter alignment between assessment purpose and instrument design, broader construct coverage, and cautious, well-validated integration of automated tools. Implications for physics teachers, instrument developers, and assessment policy are discussed.