Most studies on oil palm fresh fruit bunch (FFB) grading report accuracy as the sole measure of success, which conceals weaknesses on minority classes and says nothing about field deployability. This study builds and evaluates convolutional neural network models for four-class FFB quality grading using a multi-metric framework. An experimental design was applied to 3,200 field images collected at a palm oil mill in Riau across three daylight conditions, labeled independently by three certified graders with a Fleiss kappa of 0.81, and split 70:15:15 in stratified fashion. A baseline CNN and four transfer learning architectures were compared on accuracy, macro precision, recall, F1, Cohen kappa, inference latency, and model size. EfficientNetB0 achieved the best statistical performance at 93.54 percent accuracy and 0.914 kappa, yet per-class analysis revealed the mengkal class lagging at 88.4 percent F1, and MobileNetV2 delivered the best performance-to-cost ratio at 41 milliseconds latency. Model ranking therefore changes with the chosen metric, so grading studies must report per-class metrics, kappa, and computational cost rather than aggregate accuracy alone.
Copyrights © 2026