The rise of Artificial Intelligence (AI) challenges the integrity of traditional educational assessments, positioning the oral examination as a robust method for verifying authentic understanding. However, its large-scale implementation is hindered by significant logistical challenges, including time, resources, and evaluation consistency. This research addresses a foundational limitation for using Large Language Models (LLMs) in automated oral exams: their inherent passivity and failure to handle ambiguity, which are critical for effective assessment. To address this critical gap, we established an experimental framework and created two specialized datasets to systematically evaluate prompting techniques designed to elicit proactivity. Using the GPT-4o model, our comparative analysis reveals that the optimal strategy is highly task-dependent: for ambiguity detection (CNP), a Proactive Chain-of-Thought (PCoT) zero-shot approach achieved a near-perfect 0.99 F1-Score; for generating clarification questions (CQG), the PCoT few-shot variant was most effective; and in target-guided scenarios, a simpler proactive prompt proved superior. These findings provide foundational insights for developing automated oral examiners capable of nuanced, human-like dialogue, thereby addressing the scalability issues of traditional assessment methods.
Copyrights © 2026