Conventional personality assessment generally relies on psychometric questionnaires, such as the Big Five Inventory (BFI). However, this method is susceptible to self-report bias and requires a relatively long evaluation time. On the other hand, the development of automated video-based personality analysis systems faces high computational challenges in simultaneously integrating facial spatial features and the temporal dynamics of micro-expressions. This study aims to propose a solution in the form of a hybrid spatio-temporal neural network architecture utilizing transfer learning techniques to predict personality more objectively. The proposed system integrates the Multi-Task Cascaded Convolutional Networks (MTCNN) algorithm for automatic face detection, the PolyFace model as a spatial feature extractor, and the Long Short-Term Memory (LSTM) algorithm to model temporal relationships between frames. Experiments were conducted by comparing three optimization algorithms, namely Adam, AdaGrad, and SGD, using the ChaLearn LAP 2017 dataset. The results show that the AdaGrad optimizer achieved the best generalization and prediction performance during the validation phase, with an average accuracy of 89.13%, outperforming SGD (88.34%) and Adam (88.33%). Theoretically and mathematically, the superiority of AdaGrad in this architecture is attributed to its ability to dynamically adjust the learning rate based on the accumulation of historical gradients. This mechanism makes AdaGrad considerably more stable in extracting complex spatio-temporal micro-expression features without becoming trapped in local minima. The highest performance achieved by AdaGrad was observed in the Agreeableness dimension, reaching an accuracy of 90.14%. The main contribution of this study is the development of an efficient (low-latency), precise spatio-temporal personality prediction model that avoids local minima, making it suitable for implementation in real-world automated psychological detection systems.