Berikut terjemahan yang telah disesuaikan dengan bahasa akademik yang lazim digunakan dalam jurnal internasional (Scopus/Q1): The development of Image-based Question Answering (IQA) systems has advanced rapidly alongside the evolution of Optical Character Recognition (OCR), Information Retrieval (IR), and multimodal artificial intelligence technologies. However, the integration of OCR, retrieval mechanisms, and multimodal reasoning continues to face various conceptual and methodological challenges. Furthermore, comprehensive studies synthesizing the interrelationships among these components remain limited. This study aims to map the development, research trends, challenges, and future directions of OCR- and Information Retrieval-based techniques in Image-based Question Answering systems. A Systematic Literature Review (SLR) was conducted following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. A total of 61 indexed articles were analyzed using the Theory, Context, Characteristics, and Methodology (TCCM) framework. The findings indicate that Multimodal Deep Learning, Transformer architectures, and Attention Mechanisms constitute the primary foundations of modern Question Answering systems. In addition, the integration of Knowledge Graphs, Large Language Models (LLMs), and Retrieval-Augmented Generation (RAG) has become increasingly prominent due to their ability to enhance reasoning capabilities and contextual understanding. The review also reveals a paradigm shift from traditional feature engineering–based approaches toward knowledge-driven multimodal reasoning systems. These findings provide a foundation for the development of more adaptive, reliable, and context-aware image-based question answering systems. Despite the rapid progress of IQA research, several key challenges remain, including the semantic gap, the complexity of multimodal fusion, hallucination issues, limited availability of multilingual datasets, and the lack of system interpretability.