Indonesian Journal of Electrical Engineering and Computer Science
Vol 43, No 2: August 2026

Video summarization using deep image captioning models

Qudes M. B. Aljelawy (Al Karkh University for Science)
Sarah S. Mohammed (Ibn Sina University for Medical and Pharmaceutical Sciences)
Entessar K. Hanoun (Al Karkh University for Science)



Article Info

Publish Date
01 Aug 2026

Abstract

This research presents a novel approach for video summarization by leveraging deep image captioning models. A pretrained image captioning model, namely Salesforce's bootstrapped language image pretraining (BLIP), is used to extract keyframes from a video at regular intervals and produce natural language descriptions. These textual descriptions are then filtered for non-repetition and concatenated into a coherent summary, allowing users to understand the video’s content without viewing it in full. The proposed framework aims to improve video browsing, indexing, and retrieval efficiency, particularly for big datasets. The proposed method achieves a significant reduction in redundancy by 40% compared to raw captioning sequences. Evaluation using semantic consistency checks demonstrates that the BLIP-based framework maintains high descriptive accuracy even in complex scenes, providing a scalable solution for large-scale video indexing.

Copyrights © 2026