Praveen Sundar
Thiruvalluvar University

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Cross-modal attention fusion using vision transformers for robust student attentiveness estimation Rajasekaran Mariswamy; Praveen Sundar
International Journal of Informatics and Communication Technology (IJ-ICT) Vol 15, No 3: September 2026
Publisher : Institute of Advanced Engineering and Science

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.11591/ijict.v15i3.pp935-943

Abstract

Automated student attentiveness estimation is a fundamental component of intelligent e-learning systems and adaptive classroom analytics. Traditional convolutional and recurrent architectures often struggle to model long-range temporal dependencies and complex inter-modal relationships inherent in engagement behavior. To address these limitations, this paper proposes a cross-modal attention fusion framework built upon a vision transformer (ViT) backbone for robust student attentiveness estimation. The proposed architecture leverages patch-based visual encoding through a ViT to capture global spatial dependencies, while behavioral cues such as gaze direction, head pose, and blink dynamics are embedded into a shared latent representation space. A cross-modal multi-head attention mechanism is introduced to dynamically learn interactions between visual and behavioral modalities, replacing static weighted fusion strategies. Temporal dynamics are modeled using a Transformer encoder, enabling effective long-range sequence modeling without recurrent dependencies. Experimental evaluation on a benchmark attentiveness dataset demonstrates superior performance compared to CNN–LSTM-based models, achieving improved accuracy, F1 score, and robustness under challenging lighting and occlusion conditions. Ablation studies validate the contribution of cross-modal attention and transformer-based temporal modeling. The proposed framework maintains real-time feasibility while significantly enhancing discriminative capability.