Raffi Ardhi Naufal
Universitas Pendidikan Indonesia

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Fine-Grained Classroom Activity Recognition via Detection, Head-Pose, and Appearance Fusion Yaya Wihardi; Raffi Ardhi Naufal; Meutia Jasmine Annisa Herawan; Rani Megasari
Brilliance: Research of Artificial Intelligence Vol. 6 No. 3 (2026): Brilliance: Research of Artificial Intelligence, Article Research August 2026
Publisher : Yayasan Cita Cendekiawan Al Khwarizmi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47709/brilliance.v6i3.9453

Abstract

Automatic analysis of student behavior in classroom environments is an important prerequisite for smart learning systems that move beyond coarse engagement indicators toward behavior-specific and pedagogically interpretable feedback. Fine-grained classroom activities, however, remain difficult to recognize because several target behaviors exhibit similar visual patterns, are partially occluded, or involve subtle head and upper-body motion. This study addresses five classroom activities commonly observed in authentic instructional settings: nodding, hand-raising, smartphone use, head-supporting, and looking downward. We propose a student-centered framework that first constructs person-consistent 16-frame tracklets and then integrates three complementary sources of evidence: RGB appearance from person crops, head-pose estimation, and an explicit smartphone-related cue. This design addresses two major ambiguity clusters in classroom scenes, namely downward-looking versus smartphone use and downward-looking versus head-supporting, while preserving temporal sensitivity for nodding. The dataset comprises labeled student-centered clips extracted from multi-view classroom recordings. Evaluation follows a subject-separated and session-separated protocol using different acquisition sessions, classroom settings, and participant groups. Under this protocol, the proposed framework achieves 91.13% accuracy and 91.14% macro F1. Compared with a visual-only baseline, the full fusion model improves macro F1 by 6.21 percentage points, while outperforming the strongest non-fusion baseline by 2.83 percentage points. The confusion analysis further indicates that head-pose information and explicit smartphone cues effectively separate visually adjacent downward-oriented behaviors. These findings support multi-cue fusion as an effective strategy for fine-grained classroom activity recognition and classroom analytics under the present held-out evaluation protocol, while broader deployment and cross-session generalization should remain bounded by the current evidence.