Khin Hnin Naing
University of Medical Technology

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Near-Trivial Separability and Statistical Power Limitations in a Small Anthrax Symptom Dataset Khin Hnin Naing
Journal of Systems Engineering and Information Technology (JOSEIT) Vol. 3 No. 3 (2024)
Publisher : Ikatan Ahli Informatika Indonesia Nusantara

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29207/6n7m2c61

Abstract

A recently released IEEE DataPort dataset frames 49 binary-labeled livestock symptom records as a resource for few-shot learning and robust classification under data scarcity. We audit the dataset's suitability for this role by computing per-symptom discriminative statistics and then assessing whether n=49 leaves meaningful room for the complex models the dataset is intended to evaluate. A single symptom – the presence of bloody discharge – reconstructs the anthrax label with 91.8% accuracy (sensitivity 77.8%, specificity 100.0%, MCC φ=0.83, 95% Wilson CI [80.8%, 96.8%]), and adding a second symptom (edema) raises this to 93.9%. A leave-one-out cross-validated logistic regression and a 300-tree random forest achieve 89.8% and 93.9% accuracy respectively, numerically no better than the trivial single-symptom rule; paired McNemar tests confirm that neither classifier's errors differ significantly from those of the trivial rule (p=1.00 for both). A power analysis shows that n=49 provides only 9.1% power to detect a 2.2-percentage-point improvement over the trivial rule at α=0.05. Reliably detecting even a 6.2-point improvement would require n≈90. These results indicate that this dataset cannot discriminate between trivial threshold rules and genuinely learned classifiers, and any reported accuracy near 92–94% on this dataset should be interpreted as evidence of near-trivial separability, not few-shot learning capability.