Maize Lethal Necrosis (MLN) and Maize Streak Virus (MSV) threaten food security in Zimbabwe, yet traditional image-based deep learning models lack adaptability to varying agro-environmental conditions. This study presents a multimodal contextual learning framework that fuses visual symptoms with environmental data to improve diagnostic accuracy and interpretability. A multimodal Convolutional Neural Network integrates leaf imagery with six environmental parameters (temperature, humidity, rainfall, soil moisture, leafhopper count, days after planting) using a late-fusion architecture with multi-task learning for simultaneous disease classification and five-stage severity estimation, enhanced by Gradient-weighted Class Activation Mapping (Grad-CAM) for visual explanation. A synthesized dataset (n=9,356) representative of Zimbabwean conditions was used. The multimodal approach achieved 94.3% accuracy versus 91.5% for image-only baselines (p<0.001, 95% CI), a 33.2% relative error reduction, with MSV-MLN confusion decreasing by 40.7% (113 to 67 misclassifications). Multi-task performance reached 94.5% for classification and 81.5% for severity estimation (F1=0.805). Grad-CAM analysis revealed environmental integration enhanced attention by 38.1% and localization by 45.2%, with leafhopper count showing the strongest correlation (r=0.62). Ablation studies confirmed that environmental features provided the largest accuracy gain (+3.71%, p<0.001). This work demonstrates that integrating environmental context with visual symptoms enhances diagnostic accuracy and model interpretability, establishing a foundational proof-of-concept for context-aware agricultural AI.
Copyrights © 2026