More and more generative AI systems which includes large language models, diffusion models, and multimodal foundations are being integrated into crucial computing infrastructure, including cloud orchestration, code synthesis pipeline, healthcare decision support, and financial risk assessment. Consequently, there is greater demand for frameworks that can evaluate, guarantee, and regulate the trustworthiness of these systems. This article reviewed the development of trustworthy AI research from 2015 to 2025, and the evidence generated across four primary areas: safety and alignment, robustness and reliability, evaluation, and governance. We delivered distinctive comparative assessments of safety benchmarks, alignment methodologies (RLHF, RLAIF, DPO, Constitutional AI), and formal governance frameworks worldwide, pinpointing the critical discrepancies between regulated objectives and actual technical capability. A key finding is the Evaluation Paradox: The benchmarks most commonly relied on to certify systems as “AI safe” are, in fact, the systems least robust to distributional shift and adversarial manipulation. There is an institutional misalignment between the speed of generative AI deployment and the maturity of the governance mechanisms proposed to regulate it. We documented seven priority research challenges for the field. Researchers, system engineers, policymakers, and practitioners pursuing an evidence-based understanding of the current state-of-trustworthiness will benefit from this review.
Copyrights © 2026