Claim Missing Document
Check
Articles

Found 1 Documents
Search

Hand Gesture Recognition in Augmented Reality using Deep Learning Models Sneha Saini
International Journal of Advanced Science and Computer Applications Vol. 5 No. 1 (2026): March 2026
Publisher : Utan Kayu Publishins

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

We present a comprehensive exploration of hand gesture recognition models leveraging various deep learning architectures, including Convolutional Neural Networks (CNN), Long Short-Term Memory networks (LSTM), and pre-trained architectures such as VGG16 and ResNet. The primary objective of this research is to enhance the accuracy, robustness, and generalization capabilities of hand gesture recognition systems for applications in augmented reality (AR), human-computer interaction (HCI). We trained the models on a large-scale dataset consisting of 14,000 images representing various hand gestures, which ensured diverse and comprehensive training data. The models were designed to capture both the spatial and temporal patterns inherent in hand gestures. Additionally, pre-trained architectures like VGG16 and ResNet were employed using transfer learning techniques, which enabled these models to take advantage of their deep feature extraction capabilities. Both VGG16 and ResNet architectures were fine-tuned to adapt their learned features to the specific requirements of the hand gesture recognition task. Our experimental results demonstrate that while the CNN-LSTM models are capable of accurately recognizing gestures, the pre-trained architectures, especially ResNet, outshine them in terms of performance metrics and computational efficiency. The VGG16 model achieved the highest accuracy of 97.5% , compared to 96% for ResNet and 93% for the CNN-LSTM model. Our findings contribute to the ongoing development of more efficient and accurate hand gesture recognition systems. The insights gained from this research can be extended to future studies that explore hybrid models combining the strengths of CNN-LSTM, and pre-trained architectures to achieve even greater recognition accuracy in more challenging environments.