Accurate multiclass classification of dermoscopic skin lesions remains challenging because of high inter-class visual similarity, substantial intra-class variability, and frequent acquisition artifacts (black borders, hair occlusions, noise). We propose a unified, reproducible framework that systematically coordinates four stages: (i) artifact-aware preprocessing (field-of-view circular cropping, hair removal, CLAHE, bilateral filtering); (ii) lesion-focused segmentation via GrabCut-refined fusion and a U-Net with EfficientNet-B3 encoder; (iii) compact deep-feature extraction (EfficientNet-B7) refined by principal component analysis and Neural Spline Flow density calibration; and (iv) robust machine-learning classification. The HAM10000 dataset (n = 10,015, seven diagnostic classes) was partitioned once by stratified random sampling into training (70 %, n = 7010), validation (15 %, n = 1502), and test (15 %, n = 1503) subsets under a strictly sequential anti-leakage protocol with patient-level isolation; the test set was sequestered until terminal evaluation. External generalization was assessed on an independent ISIC 2019 subset (n = 350, 50 per class) without retraining. On the held-out HAM10000 test set, XGBoost achieved the highest accuracy of 99.47 % with an F1-score of 98.99 %, followed by LightGBM (98.20 %) and MLP (97.67 %). Ablation analysis confirmed incremental gains of +2.55 % (preprocessing), +1.75 % (segmentation), and +1.32 % (Neural Spline Flow refinement). On the external ISIC 2019 data, MLP attained the best cross-domain accuracy of 95.43 %, demonstrating that the feature backbone generalizes beyond the training distribution. The demonstrated synergy of artifact suppression, lesion-centered segmentation, and density-calibrated feature learning yields highly discriminative and generalizable representations, providing a robust foundation for reliable computer-aided dermatologic screening
Copyrights © 2026