Sofia Yosse
Politeknik Negeri Padang

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Quantifying Target Leakage in Rule-Derived Mental Health Screening Labels Rozi Meri; Ikhsan; Dian Eka Putra; Sofia Yosse; Ismael
Journal of Systems Engineering and Information Technology (JOSEIT) Vol. 3 No. 3 (2024)
Publisher : Ikatan Ahli Informatika Indonesia Nusantara

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29207/joseit.v3i3.8470

Abstract

Publicly hosted tabular datasets increasingly package screening-questionnaire items alongside computed, not clinically observed, risk and severity labels. This creates a risk of target leakage: if a label is a deterministic function of the same features given to a classifier, reported predictive performance measures rule reconstruction, not screening validity. We audit a recent IEEE DataPort dataset (n = 2,005) scoring 27 Likert items across five conditions – ADHD, Autism Spectrum Disorder (ASD), Social Pragmatic Communication Disorder (SPCD), Depression, and Anxiety – each with a binary risk flag and a three-level severity label. The dataset's documentation discloses that every label is generated from a threshold or tertile rule on the summed items of its own condition, yet its usage instructions recommend fitting Random Forest or XGBoost with SHAP directly on the 27 raw items. We verify the documented rule empirically, then compare a trivial single-feature decision stump (the summed score) against a 300-tree Random Forest trained on all 27 items, under identical 5-fold stratified cross-validation, across all ten condition-by-task combinations (5 conditions × {risk, severity}). The single-feature stump matches or exceeds the 27-feature ensemble in every combination, reaching 100% cross-validated accuracy on nine of ten targets. Random Forest feature importance further shows 72–88% of predictive weight concentrated on each condition's own items. These results indicate the pipeline the dataset recommends reconstructs a disclosed arithmetic rule rather than learning a clinically meaningful pattern; any predictive-validity study on this dataset should exclude a condition's own summed items from that condition's feature set.