International Journal of Advances in Intelligent Informatics
Vol 12, No 3 (2026): August 2026

A comprehensive review of CNN and transformer based visual feature extraction for automatic video captioning

Hemel Sharker Akash (Graduate Research Assistant, Faculty of Engineering & Technology (FET) , Multimedia University, Melaka)
Joseph Emerson Raja (Assistant Professor, FACULTY OF ENGINEERING AND TECHNOLOGY (FET), Multimedia University)
Md. Jakir Hossein (Associate Professor, FACULTY OF ENGINEERING AND TECHNOLOGY (FET), Multimedia University)



Article Info

Publish Date
31 Aug 2026

Abstract

Automatic Video Captions are increasingly important to use in applications in the accessibility, search, education, and content moderation areas. Generating accurate and context-aware captions is difficult for lengthy and complex videos. Recent surveys focus more on language models than on the vision side. This means that the contribution of visual feature extraction towards the performance of the model is not studied properly. Here we attempt to fill this gap by conducting reviews of video captioning literature from the vision viewpoint, comparing CNNs-based and transformer-based feature extraction, and detecting trends, pros and cons. Approximately 100 papers published between 2017 - 2025 were reviewed and analyzed based on base vision, dataset and evaluation metrics. The review classified the selected papers into CNN-based pipelines and transformer-based techniques. Moreover, it compared their rank measured by application popularity, dataset usage, and performance scores across popular datasets/benchmarks. The review shows that CNNs were dominated before 2021, but transformer-based approaches now lead due to their ability to capture long-range temporal dependencies and multimodal interactions. However, CNNs continue to perform well for short videos. As for Dataset usage, patterns show that ActivityNet Captions and YouCook2 are suited for long-form reasoning because of timely manner captioning, while MSR-VTT and MSVD are effective for short clips (one sentence caption). In terms of Special semantic quality metrics, CIDEr and METEOR are more semantically appropriate compared to n-gram quality metrics like BLEU. This review emphasizes transformer-based architecture, particularly multimodal models, which will be the future of video captioning. However, challenges remain in handling lengthy videos, adapting across domains and languages, and ensuring efficient deployment (small, optimized model). Addressing these gaps will add value by making video captioning systems more reliable, interpretable, and broadly usable in real-world contexts.

Copyrights © 2026






Journal Info

Abbrev

IJAIN

Publisher

Subject

Computer Science & IT

Description

International journal of advances in intelligent informatics (IJAIN) e-ISSN: 2442-6571 is a peer reviewed open-access journal published three times a year in English-language, provides scientists and engineers throughout the world for the exchange and dissemination of theoretical and ...