On BreastMNIST+ 224, we tested visual and class-prompt models with calibration, conformal sets, routing, perturbations, and evidence cards. The 546/78/156 train/validation/test artifact contained one conflicting duplicate. The ensemble achieved AUROC/AUPRC 0.885/0.738, balanced accuracy 0.842, sensitivity 0.833, and specificity 0.851. Temperature scaling reduced Brier/ECE 0.120/0.093-to-0.116/0.079. Label-conditional 90% conformal prediction achieved overall/malignant coverage 0.897/0.952 and size 1.359. Rank-2 prompts reached AUROC 0.790; blending raised sensitivity to 0.905 but lowered specificity to 0.763. Rotation and contrast degraded most. 156 cards passed checks. Results support routing, not clinical validity.
Copyrights © 2026