This study presents an indicator-level secondary analysis of de-identified pretest-posttest data from 40 Grade 10 students (20 experimental and 20 control) who studied virus concepts at two madrasahs. The experimental group received instruction incorporating three-dimensional virus models, whereas the control group received lecture-based instruction. Four constructed-response items assessed interpretation, analysis, evaluation, and inference on a 1-4 scale. Raw gains and individual normalized gains were summarized descriptively. Because gain scores were non-normally distributed, two-sided Mann-Whitney U tests served as the primary between-group comparisons; Hedges' g with bootstrap 95% confidence intervals and pretest-adjusted ANCOVA were used as complementary analyses. The experimental group showed larger mean raw gains for interpretation (0.35 vs. 0.25), analysis (0.60 vs. 0.10), evaluation (0.35 vs. 0.20), inference (0.75 vs. 0.35), and the total score (2.05 vs. 0.90). Statistically significant between-group differences were observed for analysis (U = 271.0, p = .036; g = 0.74, 95% CI [0.16, 1.40]) and the total score (U = 320.5, p < .001; g = 1.30, 95% CI [0.78, 1.97]). Inference showed the largest descriptive improvement and a moderate effect (g = 0.56), although the between-group difference was not statistically significant (p = .136). ANCOVA yielded the same substantive pattern, with significant adjusted differences for analysis and the total score but not for interpretation, evaluation, or inference. Within the constraints of a small, non-randomized, two-school design, the findings suggest that model-supported instruction was most consistently associated with analytical reasoning, whereas interpretation and evaluation may require more explicit scaffolding, argumentation, and evidence-evaluation tasks.