Huong-Giang Doan
Electric Power University

Published : 3 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 3 Documents
Search

End-to-end multiple modals deep learning system for hand posture recognition Huong-Giang Doan; Ngoc-Trung Nguyen
Indonesian Journal of Electrical Engineering and Computer Science Vol 27, No 1: July 2022
Publisher : Institute of Advanced Engineering and Science

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.11591/ijeecs.v27.i1.pp214-221

Abstract

Multi-modal or multi-view dataset that was captured from various resources (e.g. RGB and Depth) of a subject at the same time. Combination between different cues has still faced to many challenges as unique data and complementary in-formation. In adition, the proposed method for multiple modalities recognition consists of discrete blocks, such as: extract features for separative data flows, combine of features, and classify gestures. To address the challenges, we pro-posed two novel end-to-end hand posture recognition frameworks, which are integrated all steps into a convolution neuronal network (CNN) system from capturing various types of cues (RGB and Depth images) to classify hand ges-ture labels. Both frameworks use the Resnet50 backbone that was pretrained by ImageNet dataset. We proposed a novel end-to-end multi-modal frameworks, which are named attention convolution module (ACM) and gated concatenation module (GCM). Both of them are deployed, evaluated and compared on vari-ous multi-modalities hand posture datasets. Experimental results show that our proposed method outperforms with others state-of-the-art techniques (SOTA) methods.
New blender-based augmentation method with quantitative evaluation of CNNs for hand gesture recognition Huong-Giang Doan; Ngoc-Trung Nguyen
Indonesian Journal of Electrical Engineering and Computer Science Vol 30, No 2: May 2023
Publisher : Institute of Advanced Engineering and Science

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.11591/ijeecs.v30.i2.pp796-806

Abstract

In this study, we extensively analyze and evaluate the performance of recent deep neural networks (DNNs) for hand gesture recognition and static gestures in particular. To this end, we captured an unconstrained hand dataset with complex appearances, shapes, scales, backgrounds, and viewpoints. We then deployed some new trending convolution neuron networks (CNNs) for gesture classification. We arrived at three major conclusions: i) DenseNet121 architecture is the best recognition rate through almost evaluated red, green, blue (RGB) and augmentation datasets. Its performance is outstanding in most original works; ii) blender-based augmentation help to significantly increase 9% of accuracy, compared to the use of a RGB cues; iii) most CNNs can achieve impressive results at 97% accuracy when the training and testing datasets come from the same lab-based or constrained environment. Their performance is drastically reduced when dealing with gestures collected in unconstrained environments. In particular, we validated the best CNN on a new unconstrained dataset. We observed a significant reduction with an accuracy of only 74.55%. This performance can be improved up to 80.59% by strategies such as blender-based and/or GAN-based data augmentations to obtain an acceptable result of 83.17%. These findings contribute crucial factors and make fruitful recommendations for the development of a robust hand-based interface in practice
TMA-Net: a transformer-based multi-modal attention network for abnormal behavior detection Huong-Giang Doan; Ngoc-Trung Nguyen
IAES International Journal of Artificial Intelligence (IJ-AI) Vol 15, No 2: April 2026
Publisher : Institute of Advanced Engineering and Science

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.11591/ijai.v15.i2.pp1441-1450

Abstract

Abnormal behavior detection in crowded environments remains challenging due to complex motion patterns, occlusions, and domain variability. This paper presents transformer-based multi-modal attention network (TMA-Net), a unified framework that integrates red, green, and blue (RGB), optical flow (OF), and heat map (HM) modalities through a dual-stage attention fusion mechanism. The system employs you only look once version 11 (YOLOv11) for human localization and vision transformer (ViT)-B/16 for feature encoding, followed by intra-modal self-attention and cross-modal fusion to capture fine-grained spatial–temporal and motion energy dependencies. Extensive experiments on six public benchmarks as UMN, Crowd-11, UBNormal, ShanghaiTech, CUHK Avenue, UCSD Ped2, and EPUAbN dataset, demonstrate that TMA-Net achieves up to 97.5% area under the curve (AUC) and 96–100% accuracy, outperforming previous other state-of-the-art approaches. These results highlight the framework’s strong generalization and robustness across both single- and cross-dataset evaluations, underscoring its potential for reliable deployment in real intelligent surveillance systems.