The proliferation of generative artificial intelligence (generative AI) tools has shaken the fundamental assumptions of learning evaluation in higher education, particularly the premise that a piece of writing submitted by a student reliably reflects that individual's own competence. This study employs a qualitative design with a library research approach to examine how the field of learning evaluation is responding to this disruption, focusing on three areas: (1) the redesign of assessment instruments toward AI-resistant authentic tasks; (2) the validity and reliability of AI detection instruments used to safeguard evaluation integrity; and (3) the paradigm shift from summative toward formative evaluation enabled by generative AI-based feedback. Data were collected from indexed journals, one international-scale multi-tool empirical study, and classical evaluation theory literature published between 2011 and 2025, and were subsequently analyzed thematically. The findings indicate that the assessment redesign literature converges on process-oriented and higher-order thinking tasks that are structurally difficult to fully automate, in line with constructive alignment theory. However, empirical testing of AI detection tools reveals serious reliability issues, with accuracy rates below 80% for the majority of tools tested, as well as documented risks of false accusation, rendering such tools unsuitable as the sole evidence in high-stakes evaluation decisions. Meanwhile, automated feedback based on generative AI shows potential for scaling rapid formative feedback, although its pedagogical quality remains variable and continues to require human oversight. This study concludes that an accountable evaluation system in the era of generative AI must simultaneously redesign the object being assessed, treat detection technology as supporting rather than definitive evidence, and redirect the culture of evaluation from summative gatekeeping toward sustainable and dialogic formative practice.
Copyrights © 2026