International Journal of Engineering, Science and Information Technology
Vol 6, No 3 (2026)

From Batch Prediction to Self-Healing ML Systems: A Production MLOps Framework for Autonomous Model Maintenance

Kushwanth Chowdary Kandala (Independent Researcher)



Article Info

Publish Date
28 Jul 2026

Abstract

Machine learning (ML) systems deployed in large-scale healthcare Software-as-a-Service (SaaS) environments frequently experience performance degradation after production release due to data drift, evolving data distributions, pipeline latency anomalies, and declining predictive accuracy. These issues often remain undetected until they significantly affect operational performance, resulting in prolonged incident detection, costly engineering interventions, and increased business risk. This study proposes a self-healing machine learning framework that transforms conventional batch-oriented ML pipelines into autonomous, event-driven operational systems capable of continuously monitoring, diagnosing, and recovering from production failures. The proposed architecture integrates four complementary capabilities: statistical data drift detection, automated retraining triggers, canary-based deployment and promotion workflows, and multi-tier observability that connects model behaviour with automated operational responses. Together, these components establish a closed feedback loop that enables continuous adaptation while minimizing manual intervention. The framework is evaluated through a production case study involving a healthcare long-term care SaaS platform supporting large-scale clinical operations. Empirical results demonstrate substantial operational improvements following deployment of the self-healing architecture. Mean time to detect model drift decreased from 18.4 hours to 2.1 hours, while mean time to recovery was reduced from 9.2 hours to 1.8 hours. In addition, the number of monthly incidents requiring manual intervention declined from 34 to 6, indicating significant gains in operational resilience and engineering efficiency. The proposed framework is implemented using widely adopted open-source technologies, including MLflow, Kubeflow Pipelines, Feast, Great Expectations, and Argo Rollouts, allowing seamless integration with existing Kubernetes-based enterprise infrastructures. The findings demonstrate that autonomous, event-driven maintenance substantially improves the reliability, scalability, and maintainability of production ML systems, providing a practical engineering architecture for resilient AI operations in mission-critical healthcare environments

Copyrights © 2026