Selecting suitable crop types based on soil and environmental conditions is essential for improving agricultural productivity. Machine learning enables crop classification using soil and climate characteristics, but using all features may increase model complexity without necessarily improving performance. This study analyzes the effect of Information Gain-based feature selection on Random Forest, Decision Tree, and K-Nearest Neighbors (KNN) for crop classification. The farming.csv dataset contains 2,200 samples, seven input features—nitrogen (N), phosphorus (P), potassium (K), temperature, humidity, pH, and rainfall—and 22 crop classes. The experiment employed an 80:20 train-test split with random_state 42. Feature selection using mutual_info_classif with a threshold >1.0 reduced the features to six: humidity, K, rainfall, P, temperature, and N. Model performance was evaluated using accuracy, precision, recall, and F1-score. Using all features, Random Forest, Decision Tree, and KNN achieved accuracies of 99.32%, 98.64%, and 97.05%, respectively. After feature selection, Random Forest and Decision Tree achieved 99.09%, while KNN remained at 97.05%. These results indicate that Information Gain can reduce the feature set by one feature while maintaining nearly unchanged classification performance.
Copyrights © 2026