The study aimed to develop an accurate predictive model to identify individuals without health insurance due to unemployment. The study participants comprised 2,496 individuals who lacked general health insurance coverage in the United States for the period 2020-2023. Microdata records from the National Health Interview Survey available in the Integrated Public Use Microdata Series (IPUMS) served as the data source. The response variable was a binary variable indicating whether an individual had no health insurance (uninsured) due to unemployment. Statistical analyses involved descriptive measures: frequencies, percentages, and quartiles and inferential tests: Pearson’s chi-squared test, Wilcoxon rank-sum test, and Goodman and Kruskal’s measures. For predictive analysis, data were partitioned into training and testing sets, with class imbalance addressed using the Synthetic Minority Over-sampling Technique (SMOTE). Binary logistic regression and probit models were applied, and model performance was assessed using accuracy, precision, recall, and F1 score. Most participants reported being in good, very good, or excellent health. The demographic and health-related factors in this study had limited standalone predictive values. However, Probit and Logit models identified obesity, age, lower education, and lack of healthcare access as critical risk factors. Furthermore, the prediction performance metrics for both Logit and Probit models were identical.
Copyrights © 2026