This systematic literature review synthesizes empirical and conceptual scholarship on the application of AI-in history education between 2020 and 2026, interrogating research trends, pedagogical affordances, ethical risks, and policy implications. Employing a PRISMA-guided search across SCOPUS, Web of Science, and ERIC and a PICOC-framed analytic scope, the study screened 312 records (after deduplication) and conducted in-depth analysis of 10 peer-reviewed articles using thematic synthesis, bibliometric mapping, and quality appraisal (CASP). Findings indicate a clear disciplinary trajectory from conceptual explorations toward classroom-embedded uses of generative AI and learning analytics, with demonstrated benefits in student engagement, adaptive feedback, dialogic inquiry, and immersive simulation design that can scaffold elements of historical thinking. However, the corpus reveals persistent limitations: small-scale designs, Western-centric datasets, epistemic misalignments with historical heuristics (sourcing, contextualization, evidentiary weighting), algorithmic bias, opacity, and equity gaps that risk historiographical flattening and the reproduction of structural inequalities. The review concludes that AI’s pedagogical value in history depends on three interlocking priorities targeted teacher professional development integrating digital and historiographical literacies, context-sensitive governance ensuring explainability and data justice, and ethical-by-design AI architectures with traceable source attribution to safeguard epistemic rigor and inclusive learning outcomes.