Source shift can make phishing detectors appear stronger while concealing unsafe class-specific errors. This study evaluated a three-class workflow (Legitimate, Suspicious, and Phishing) using source-aware governance, controlled augmentation, family-disjoint diagnostics, and pre-specified safety gates. The frozen training set contained 3,732 records from four source families. Evaluation used a locked 584-record validation set, a 587-record internal test, a 788-record SMS compound test, and a 200-record SpamAssassin family-disjoint diagnostic. Classical word/character TF-IDF logistic regression with cell-balanced augmentation achieved internal macro-F1 of 0.9633 and compound macro-F1 of 0.5461. DistilRoBERTa augmentation improved compound macro-F1 by 0.1524 versus its reference (95% CI 0.1143-0.1914), but reduced phishing recall by 0.2374 and failed the severe-error gate. A hierarchical transformer repair achieved internal macro-F1 of 0.9668 and compound macro-F1 of 0.5119; its compound gain was 0.1774 (95% CI 0.1410-0.2152), while phishing recall remained 0.2323 below the reference. Governance review included 492 dual-human consensus records, eight independently adjudicated disagreements, and protected-overlap screening. Candidate replacement sources were excluded under provenance, licensing, channel, privacy, taxonomy, or overlap criteria, so no records were added after the training freeze. The results show that source-aware augmentation can improve targeted and aggregate metrics, but aggregate gains are insufficient when phishing-recall and severe-error safety criteria fail. The study supports a governance and evaluation contribution, not global source resistance or autonomous deployment.