This study proposes a machine learning–based rice yield prediction system with a self-updating mechanism, using Gorontalo Province, Indonesia, as a case study. The system integrates daily climate data from Open-Meteo with agricultural statistics from the Central Bureau of Statistics (BPS) to support data-driven decision-making in agriculture. A key challenge addressed in this study is the limited availability of yield data, which are provided only at an annual scale for the period 2018–2024, without seasonal labels. To overcome this limitation, a temporal disaggregation approach is adopted to construct initial seasonal yield labels (M1, M2, M3). These constructed labels serve as approximations, enabling the development of a seasonal prediction model under data-constrained conditions. Several machine learning algorithms, namely Gradient Boosting, Random Forest, XGBoost, Ridge Regression, and Linear Regression, are evaluated using Leave-One-Out Cross-Validation (LOO-CV). Model performance is assessed using Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and the coefficient of determination (R²). The results indicate that ensemble-based models outperform linear baselines, with Gradient Boosting providing the best balance between prediction accuracy and model stability. The main contribution of this study is the design of an adaptive prediction system with a self-updating mechanism that supports periodic retraining and dynamic model evaluation. At its current stage, this mechanism is positioned as an initial framework rather than a fully validated continuous learning system. The proposed system is implemented as a web-based platform that supports yield prediction and planting season recommendations, providing a scalable foundation for intelligent agricultural systems in data-limited environments.
Copyrights © 2026